One source for head: SiteHead, JSON-LD escaping and hreflang

All three HTML shells render one component that emits every head tag, with the data assembled on the server. Plus a detail worth remembering: angle brackets in JSON-LD must be escaped.

The easiest place for a site to grow crooked is the head. One og:image on the home page, a JSON-LD block on posts, half of it copied into the 404 page, and three months later nobody can say which page carries which tags.

A single source

components/SiteHead.astro emits every head tag and shellHead() in server/shell.ts assembles the data. All three HTML shells render it: the two layouts and the language entry page. SEO changes touch exactly those two files.

Escaping inside JSON-LD

Structured data is inline JSON, so after serialising, every < must become \u003c:

const json = JSON.stringify(data).replace(/</g, '\\u003c');

Without it, a post title containing a closing script tag closes the element early and the rest of the HTML is parsed as script. This is not hypothetical: titles are author-controlled input.

hreflang and canonical

Each page emits three hreflang links plus x-default, which always points at Chinese. canonical only ever names the apex-with-www origin, because the bare domain 301s at the Nginx layer and serving both hosts would be duplicate content.

robots.txt is a build artifact too

It is generated by pages/robots.txt.ts and references FALLBACK_ORIGIN from lib/site.ts rather than sitting in public/. The reason is practical: a static file gets forgotten during a domain migration, a generated one cannot be.

Why be this strict

Head mistakes have no visual feedback. A wrong og:image leaves the page looking perfectly normal and only breaks the share card; a wrong canonical is visible only to a crawler. Treating single-sourcing as a hard constraint is the only way this class of bug gets checked statically.

← Back to all posts

Comments

…