1. Repository and ownership boundary
TechAPI/data is the canonical store for smartphone, SoC, GPU, CPU and brand JSON. Validation, loading and delivery live in TechEngine so data licensing and release cadence remain independent from executable code. Each repository pins the other by submodule SHA, making the engine-data pair reproducible.
2. Contribution and validation pipeline
A dependency-free validator first checks directory rules, slugs, required values and source_urls on a data pull request. TechEngine then applies range, duplicate, reference-integrity and cross-field rules. Failures report the record path and reason in CI; only changes that pass every gate move from develop to main.
3. Runtime read paths
The TechEngine loader reads the pinned TechAPI checkout into common domain models. One path maps records to SQLModel entities, seeds SQLite or PostgreSQL and serves FastAPI queries. A second path emits a versioned static JSON tree and coverage metadata from the same validated snapshot. Database and dump delivery never create competing data sources.
4. Consumer and deployment boundary
FastAPI serves filtered queries, while static dumps support CDNs, application builds and the Astro explorer without a running server. The documentation site stays available during API downtime, and downstream products such as TechPicks can pin a curated build catalog. Public read delivery is deliberately kept separate from any authenticated write surface.
5. Refresh and failure behavior
Scheduled workflows collect upstream material into a staged change and open a dated pull request only after the same integrity gates pass. Actor, commit-marker and no-change checks prevent submodule webhook loops. When an upstream format changes, collection is partial or schema validation fails, the last public snapshot remains intact and the pipeline stops at a reviewable pull request.
6. Data contract and ownership
Category JSON is the source of truth while _verify ledger, caches and status own verification history. The Astro site is a read-only consumer and does not own API-server state, allowing data and presentation changes to ship independently.
Schema, URL and cross-reference checks that traverse the full tree grow with the dataset. Cache hit rate, changed-file incremental validation and static v1 artifact size need measurement; throughput and P95 build time have not been measured.
8. Security, observability and debt
CODEOWNERS and PR gates make supply-chain changes reviewable, while long-term source-URL validity and submodule revisions remain separate risks. Failed workflows and the ledger provide audit evidence, but availability SLOs and public-endpoint monitoring are missing.