Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .github/actions/test/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,14 @@ runs:
- name: No executable scripts in the markup
shell: bash
run: python3 tools/check_scripts.py dist
# The host the site claims to be is a property of the build output, not of
# any one source file: the layout derives canonical/og:url/JSON-LD from
# `site` in astro.config.mjs while robots.txt and sitemap.xml are copied
# verbatim from public/. The gate reads dist/ and fails on any claim that
# is not https://www.termca.de, and on the old host appearing anywhere.
- name: Every host claim is www.termca.de
shell: bash
run: python3 tools/check_domain.py dist
# A page that builds can still be broken in a browser. This loads the
# built site in headless Chrome -- preinstalled on the runner, so there is
# no browser download -- and fails on console errors, uncaught exceptions,
Expand Down
16 changes: 13 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -155,6 +155,16 @@ line up.

## Domain

`src/layouts/Full.astro`, `public/robots.txt` and `public/sitemap.xml` assume
`https://termcade.com`. If the site lands somewhere else, those three files are
the only places the host appears.
The marketing domain is `termca.de`, deployed on Vercel. The apex
`https://termca.de` redirects with HTTP 308 to `https://www.termca.de/` — a
redirect configured in the Vercel project's Domains settings, not in this
repository — so the canonical technical URL is `https://www.termca.de/`. The
application and registry are separate hosts (`app.termca.de`, `api.termca.de`)
and are not touched by anything here.

The host is set once, as `site` in `astro.config.mjs`; the canonical link,
`og:url` and the SoftwareApplication JSON-LD in `src/layouts/Full.astro` all
derive from `Astro.site`. `public/robots.txt` and `public/sitemap.xml` are
copied as-is, so they carry the host literally. `tools/check_domain.py` checks
every one of those claims in the built output — and that no `termcade.com`
survives anywhere in `dist/` — and runs in CI.
10 changes: 9 additions & 1 deletion astro.config.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -4,5 +4,13 @@ import { defineConfig } from 'astro/config';
// No integrations. The site is one static page that ships no JavaScript, so
// there is no framework runtime to add and nothing to hydrate. The arcade and
// its registry are the application; this is only the front door.
//
// `site` is the one place the host is configured. The apex termca.de redirects
// 308 to www in the Vercel project settings, so the canonical URL is www.
// Everything on the page that names the host — canonical link, og:url, and
// the SoftwareApplication JSON-LD — derives from Astro.site; robots.txt and
// sitemap.xml live in public/, are copied as-is, and carry the host literally.
// https://astro.build/config
export default defineConfig({});
export default defineConfig({
site: 'https://www.termca.de',
});
2 changes: 1 addition & 1 deletion public/robots.txt
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
User-agent: *
Allow: /

Sitemap: https://termcade.com/sitemap.xml
Sitemap: https://www.termca.de/sitemap.xml
2 changes: 1 addition & 1 deletion public/sitemap.xml
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,6 @@
moves on every unrelated deploy is a worse signal than none. -->
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://termcade.com/</loc>
<loc>https://www.termca.de/</loc>
</url>
</urlset>
8 changes: 6 additions & 2 deletions src/layouts/Full.astro
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,11 @@ const {
'A terminal arcade in Go. Classic games drawn at sub-cell resolution with block pixels, shipped with Asteroid and Tetris, running every game as a sandboxed WebAssembly package — and the same binary is the dev kit for writing your own.',
} = Astro.props;

const canonical = new URL(Astro.url.pathname, 'https://termcade.com');
// The host appears once, as `site` in astro.config.mjs; everything here
// derives from it. Astro.site is that URL parsed, so a config change is the
// only change a move ever needs.
const site = Astro.site ?? new URL('https://www.termca.de');
const canonical = new URL(Astro.url.pathname, site);

// Two entries rather than one. The SoftwareApplication is the thing people are
// being sent to install; the FAQPage is the questions this page actually
Expand All @@ -24,7 +28,7 @@ const structuredData = [
'@context': 'https://schema.org',
'@type': 'SoftwareApplication',
name: 'termcade',
url: 'https://termcade.com',
url: site.origin,
applicationCategory: 'GameApplication',
operatingSystem: 'Linux, macOS, Windows',
description,
Expand Down
152 changes: 152 additions & 0 deletions tools/check_domain.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""Fail if the built site names any host other than https://www.termca.de.

The marketing domain is termca.de: the apex redirects 308 to www in the Vercel
project settings, so every host claim the build emits must be www.termca.de.
The claims are spread across three files with two mechanisms — the layout
derives canonical/og:url/JSON-LD from `site` in astro.config.mjs, while
robots.txt and sitemap.xml are public/ files copied verbatim — so no single
source file is a complete picture. This gate reads the built output, which is
the only place all of them land together.

Checks, on dist/:

- index.html: exactly one <link rel="canonical"> and one og:url, both
https://www.termca.de/; the SoftwareApplication entry in the JSON-LD has
url https://www.termca.de
- robots.txt: every Sitemap: value — all of them parsed, not just the first
— is https://www.termca.de/sitemap.xml
- sitemap.xml: every <loc> value — the file is parsed as XML, not grepped —
is https://www.termca.de/
- no file anywhere in dist/ contains the old host termcade.com

The sitemap and robots checks validate every value rather than searching for
the right one: a build that kept the canonical claim but added a second,
foreign one must still fail.

python3 tools/check_domain.py dist
"""
import json
import pathlib
import sys
import xml.etree.ElementTree as ET
from html.parser import HTMLParser

HOST = 'https://www.termca.de'
OLD_HOST = 'termcade.com'
SITEMAP_NS = '{http://www.sitemaps.org/schemas/sitemap/0.9}'


class Head(HTMLParser):
def __init__(self):
super().__init__()
self.canonical = []
self.og_url = []
self.ld = []
self._in_ld = False

def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
if tag == 'link' and attrs.get('rel') == 'canonical':
self.canonical.append(attrs.get('href', ''))
elif tag == 'meta' and attrs.get('property') == 'og:url':
self.og_url.append(attrs.get('content', ''))
elif tag == 'script' and attrs.get('type', '').lower() == 'application/ld+json':
self._in_ld = True

def handle_endtag(self, tag):
if tag == 'script':
self._in_ld = False

def handle_data(self, data):
if self._in_ld:
self.ld.append(data)


def software_application_urls(head):
urls = []
for blob in head.ld:
data = json.loads(blob)
entries = data if isinstance(data, list) else [data]
for entry in entries:
if isinstance(entry, dict) and entry.get('@type') == 'SoftwareApplication':
urls.append(entry.get('url', ''))
return urls


def robots_sitemaps(path):
"""Every Sitemap: directive in robots.txt. A missing file is one failure,
not a silently empty list."""
if not path.is_file():
return None
return [
value.strip()
for line in path.read_text().splitlines()
for directive, _, value in [line.partition(':')]
if directive.strip().lower() == 'sitemap'
]


def sitemap_locs(path):
"""Every <loc> in the sitemap, parsed as XML. Returns None for a missing
file; malformed XML exits — a sitemap that is not XML is broken however
it names the host."""
if not path.is_file():
return None
try:
root = ET.parse(path).getroot()
except ET.ParseError as error:
sys.exit(f'{path}: not valid XML ({error}) — a broken sitemap is a '
f'broken host claim')
return [(loc.text or '').strip() for loc in root.iter(f'{SITEMAP_NS}loc')]


def main(root):
bad = []
root = pathlib.Path(root)
index = root / 'index.html'
if not index.is_file():
sys.exit(f'{root}: no index.html — was the site built?')

head = Head()
head.feed(index.read_text())

for what, got, want in [
('<link rel="canonical">', head.canonical, [f'{HOST}/']),
('og:url', head.og_url, [f'{HOST}/']),
('SoftwareApplication JSON-LD url', software_application_urls(head), [HOST]),
]:
if got != want:
bad.append(f'{index}: {what} is {got or ["<missing>"]}, expected {want}')

for name, claim, got, want in [
('robots.txt', 'Sitemap:', robots_sitemaps(root / 'robots.txt'),
[f'{HOST}/sitemap.xml']),
('sitemap.xml', '<loc>', sitemap_locs(root / 'sitemap.xml'),
[f'{HOST}/']),
]:
path = root / name
if got is None:
bad.append(f'{path}: missing')
elif got != want:
bad.append(f'{path}: {claim} values are {got or ["<none>"]}, '
f'expected exactly {want}')

for path in sorted(root.rglob('*')):
if path.is_file() and OLD_HOST in path.read_bytes().decode('utf-8', 'replace'):
bad.append(f'{path}: still names the old host {OLD_HOST}')

if bad:
print(f'{root}: host claims are wrong — the canonical host is {HOST}:',
file=sys.stderr)
for line in bad:
print(f' {line}', file=sys.stderr)
sys.exit(1)
print(f'{root}: canonical, og:url and JSON-LD name {HOST}; robots and '
f'sitemap point at it; no {OLD_HOST} anywhere', file=sys.stderr)


if __name__ == '__main__':
if len(sys.argv) != 2:
sys.exit('usage: check_domain.py dist')
main(sys.argv[1])