← 주요 프로젝트← Key projects

오픈 소스 프로젝트

OPEN SOURCE PROJECT

TechAPI

TechAPI

스마트폰과 프로세서의 사양을 공통 JSON 형식으로 정리하는 프로젝트입니다. 데이터를 검증하고 웹과 API로 제공합니다.

Device specifications organized in a shared JSON format, validated and published through a website and API.

TYPESCRIPTASTROJSON SCHEMAGITHUB ACTIONSVERCEL
담당 작업Work involved

JSON 스키마, 데이터 검증, API와 사이트 구현

JSON schemas, data validation, API and website implementation

현재 상태Current status

데이터셋과 사이트 공개 · 데이터 수집·검증 보완 중

Dataset and website published; collection and validation being improved

01

프로젝트 개요

TechAPI의 문제 정의, 범위와 담당 영역

TechAPI는 스마트폰, SoC, GPU, CPU와 브랜드 사양을 공통 JSON 형식으로 정리하는 프로젝트다. 제조사마다 다른 단위와 이름을 맞추고, 검증한 데이터를 웹과 API로 제공한다.

범위

  • TechAPI 저장소는 원본 데이터, 기여 규칙, 경량 검증기와 공개 Astro 사이트를 담당한다.
  • TechEngine은 전체 스키마 검증, SQLModel 적재, FastAPI 읽기 API, 정적 JSON 덤프, 커버리지 검사와 수집 자동화를 담당한다.
  • 데이터는 브랜드·SoC·스마트폰·GPU·CPU의 다섯 범주로 분리한다.

담당 영역

데이터 경로와 JSON 스키마, 검증 작업, API와 사이트 배포를 구현했다. 데이터셋과 API는 다른 프로젝트에서도 사용할 수 있게 공개했다.

현재 상태: 공개 운영 중. 완료된 기능과 다음 수집원 확장은 roadmap.md에서 구분한다.

01

Overview

Problem, scope and ownership of TechAPI

TechAPI organizes smartphone, SoC, GPU, CPU and brand specifications in a shared JSON format. It normalizes vendor names and units, validates the data, and publishes it through a website and API.

Scope

  • The TechAPI repository owns source data, contribution rules, a lightweight validator and the public Astro site.
  • TechEngine owns full validation, SQLModel seeding, the FastAPI read API, static JSON dumps, coverage checks and ingestion automation.
  • Data is separated into brand, SoC, smartphone, GPU and CPU categories.

Ownership

The work includes data paths, JSON schemas, validation jobs, and API and website deployment. The dataset and API are published for use in other projects.

Current status: publicly available. Completed and planned work are explicitly separated in roadmap.md.

02

아키텍처

데이터 저장소와 엔진, API, 정적 사이트의 연결 구조

1. 저장소와 소유권 경계

TechAPI/data가 스마트폰, SoC, GPU, CPU와 브랜드 JSON의 정본이다. 데이터 라이선스와 변경 주기를 실행 코드에서 분리하기 위해 검증·적재·서빙 로직은 TechEngine에 둔다. 두 저장소는 상대 저장소를 submodule SHA로 고정하므로 어떤 엔진 버전이 어떤 데이터 snapshot을 처리했는지 재현할 수 있다.

2. 기여와 검증 파이프라인

데이터 PR은 먼저 표준 라이브러리만 사용하는 경량 검증기를 지난다. 디렉터리 규칙, slug, 필수 값과 source_urls를 빠르게 확인한 뒤 TechEngine이 범위, 중복, 참조 무결성과 교차 필드 규칙을 검사한다. 실패는 레코드 경로와 이유를 CI에 남기고, 모든 gate를 통과한 변경만 develop에서 main으로 승격된다.

3. 런타임 읽기 경로

TechEngine loader가 고정된 TechAPI checkout을 읽어 공통 도메인 모델로 변환한다. 한 경로는 SQLModel entity로 매핑해 SQLite 또는 PostgreSQL을 seed하고 FastAPI query에 연결한다. 다른 경로는 같은 검증 결과에서 버전이 붙은 정적 JSON tree와 coverage metadata를 생성한다. DB와 정적 dump가 같은 정본을 사용해 전달 방식이 달라도 값의 출처가 갈라지지 않는다.

4. 소비자와 배포 경계

FastAPI는 필터와 조회가 필요한 소비자를 담당하고, 정적 dump는 CDN·앱 build·Astro 탐색 사이트가 서버 없이 읽는다. Astro 사이트는 API 장애와 독립적으로 배포되며 TechPicks 같은 downstream 앱은 필요한 범위만 build catalog로 고정할 수 있다. 읽기 전용 공개 데이터와 인증이 필요한 쓰기 API는 의도적으로 분리한다.

5. 자동 갱신과 실패 처리

예약 workflow가 원천을 수집해 staging 변경을 만들고 동일한 무결성 gate 뒤에서 날짜가 붙은 PR을 연다. submodule 동기화에는 actor·commit marker와 변경 없음 검사를 둬 webhook loop를 막는다. 원천 형식 변경, 부분 수집, schema 불일치가 발생하면 기존 공개 snapshot을 유지하고 검토 가능한 PR에서 멈추는 것이 기본 실패 전략이다.

6. 데이터 계약과 상태 소유권

카테고리별 JSON이 원본이고 _verify ledger·cache·status가 검증 이력을 소유한다. Astro 사이트는 읽기 전용 소비자이며 API 서버 상태를 소유하지 않는다. 이 경계를 지켜 데이터 변경과 표현 변경을 독립 배포할 수 있다.

7. 성능·확장 병목

전체 트리를 순회하는 schema·URL·교차 참조 검증은 데이터 증가에 비례해 길어진다. cache 적중률, 변경 파일 기반 증분 검증과 정적 v1 산출물 크기를 측정해야 하며 현재 처리량·P95 빌드 시간은 측정되지 않았다.

8. 보안·관측·기술 부채

CODEOWNERS와 PR gate는 공급망 변경을 검토 가능하게 만들지만 출처 URL의 장기 유효성과 submodule revision 신뢰는 별도 위험이다. 실패 workflow와 ledger는 감사 근거를 남기지만 서비스 가용성 SLO와 공개 endpoint 모니터는 미구현이다.

02

Architecture

Connection between the dataset, engine, API and static site

1. Repository and ownership boundary

TechAPI/data is the canonical store for smartphone, SoC, GPU, CPU and brand JSON. Validation, loading and delivery live in TechEngine so data licensing and release cadence remain independent from executable code. Each repository pins the other by submodule SHA, making the engine-data pair reproducible.

2. Contribution and validation pipeline

A dependency-free validator first checks directory rules, slugs, required values and source_urls on a data pull request. TechEngine then applies range, duplicate, reference-integrity and cross-field rules. Failures report the record path and reason in CI; only changes that pass every gate move from develop to main.

3. Runtime read paths

The TechEngine loader reads the pinned TechAPI checkout into common domain models. One path maps records to SQLModel entities, seeds SQLite or PostgreSQL and serves FastAPI queries. A second path emits a versioned static JSON tree and coverage metadata from the same validated snapshot. Database and dump delivery never create competing data sources.

4. Consumer and deployment boundary

FastAPI serves filtered queries, while static dumps support CDNs, application builds and the Astro explorer without a running server. The documentation site stays available during API downtime, and downstream products such as TechPicks can pin a curated build catalog. Public read delivery is deliberately kept separate from any authenticated write surface.

5. Refresh and failure behavior

Scheduled workflows collect upstream material into a staged change and open a dated pull request only after the same integrity gates pass. Actor, commit-marker and no-change checks prevent submodule webhook loops. When an upstream format changes, collection is partial or schema validation fails, the last public snapshot remains intact and the pipeline stops at a reviewable pull request.

6. Data contract and ownership

Category JSON is the source of truth while _verify ledger, caches and status own verification history. The Astro site is a read-only consumer and does not own API-server state, allowing data and presentation changes to ship independently.

7. Performance and scale bottlenecks

Schema, URL and cross-reference checks that traverse the full tree grow with the dataset. Cache hit rate, changed-file incremental validation and static v1 artifact size need measurement; throughput and P95 build time have not been measured.

8. Security, observability and debt

CODEOWNERS and PR gates make supply-chain changes reviewable, while long-term source-URL validity and submodule revisions remain separate risks. Failed workflows and the ledger provide audit evidence, but availability SLOs and public-endpoint monitoring are missing.

구조와 데이터 흐름

Structure and data flow

도식을 누르면 크게 볼 수 있습니다. 구현 여부와 참고한 코드도 표시했습니다.

Select a diagram to enlarge it. Labels show implementation status and source files.

01
시스템 컨텍스트 · 신뢰 경계System context · trust boundaries기여자·소비자·GitHub Pages 사이의 공개 데이터 신뢰 경계입니다.Public-data trust boundaries across contributors, consumers and GitHub Pages.
선택하면 전체 화면에서 세부 구조와 근거 번호를 볼 수 있습니다.Select to inspect the structure and evidence references full screen.
02
런타임 · 모듈 · 상태 소유권Runtime · modules · state ownershipPython 검증기와 정적 Astro 소비면의 소유권을 분리합니다.Separates ownership between the Python validator and static Astro consumer.
선택하면 전체 화면에서 세부 구조와 근거 번호를 볼 수 있습니다.Select to inspect the structure and evidence references full screen.
03
핵심 사용자 흐름 · 요청 시퀀스Core user flow · request sequence데이터 변경이 검증을 통과해 공개 산출물이 되는 흐름입니다.How a data change becomes a validated public artifact.
선택하면 전체 화면에서 세부 구조와 근거 번호를 볼 수 있습니다.Select to inspect the structure and evidence references full screen.
04
데이터 · 배포 · 보안 · 복구Data · delivery · security · recovery원본 JSON, 검증 상태, 배포와 실패 차단을 함께 보여줍니다.Connects source JSON, verification state, delivery and failure containment.
선택하면 전체 화면에서 세부 구조와 근거 번호를 볼 수 있습니다.Select to inspect the structure and evidence references full screen.
03

기술적 결정

TechAPI에서 선택한 경계와 그 트레이드오프

데이터와 엔진 분리

하나의 저장소가 단순하지만 코드 라이선스와 데이터 라이선스, 릴리스 주기가 결합된다. 데이터는 CC-BY-SA, 엔진은 MIT로 분리해 미러링과 재사용 경계를 명확히 했다. 대신 두 저장소의 버전을 맞추는 자동화가 필요해졌다.

API와 정적 덤프 병행

FastAPI는 검색과 필터에 유연하지만 항상 실행 중인 서버가 필요하다. 정적 덤프는 기능이 제한되는 대신 CDN에서 저렴하고 안정적으로 제공된다. 동일한 검증 데이터를 두 방식으로 배포해 소비자가 환경에 맞게 선택하도록 했다.

경로 자체를 계약으로 사용

범주·제조사·연도·segment를 디렉터리에 반영하고 slug를 kebab-case로 제한했다. 파일 이동 비용은 생기지만 리뷰 단계에서 데이터의 위치와 중복을 빠르게 판단할 수 있다.

출처 필수화

모든 레코드에 canonical source_urls를 요구한다. 기여 속도보다 검증 가능성과 추적 가능성을 우선한 결정이다.

03

Decisions

TechAPI boundaries and their tradeoffs

Separate data from the engine

A monorepo is simpler, but it couples code and data licensing and release cadence. Keeping CC-BY-SA data separate from the MIT engine makes mirroring and reuse explicit, at the cost of synchronization automation.

Offer an API and static dumps

FastAPI supports flexible queries but requires a running service. Static dumps are less expressive but cheap and reliable on a CDN. Both are generated from the same validated data so consumers can choose the appropriate delivery model.

Treat paths as part of the contract

Category, manufacturer, year and segment appear in directory paths, while slugs are restricted to kebab-case. Moves become deliberate migrations, but reviewers can identify placement and duplication without executing the engine.

Require sources

Every record must include canonical source_urls. Contribution speed is intentionally traded for traceability and reviewable evidence.

04

검증

데이터와 엔진의 자동 검증 범위와 한계

자동 검사

  • TechAPI의 python -m app.validate는 표준 라이브러리만으로 실행되어 모든 데이터 PR에서 빠르게 동작한다.
  • TechEngine은 단위·통합 테스트와 lint, type check를 별도 워크플로에서 실행한다.
  • 전체 데이터 검사는 스키마, 값 범위, slug와 참조의 유일성, 출처 누락을 확인한다.
  • 커버리지 작업은 외부 제품 목록과 현재 데이터셋의 차이를 보고서로 만든다.

재현 경로

로컬에서는 검증 → 데이터베이스 seed → FastAPI 실행 → 정적 dump 생성 순으로 같은 데이터를 확인할 수 있다. Docker Compose는 PostgreSQL 환경을 제공하고 기본 개발 경로는 SQLite를 사용한다.

알려진 한계

제조사 페이지 형식이 바뀌거나 지역별 제품명이 다르면 수집 규칙을 갱신해야 한다. 전체 데이터 정확도를 단일 백분율로 측정한 공개 결과는 없어 측정되지 않음으로 표시한다. 자동화는 출처와 구조를 검증하지만 원천 정보 자체의 오류까지 보장하지 않는다.

04

Validation

Automated checks, reproducibility and known limits

Automated checks

  • python -m app.validate uses only the Python standard library and runs quickly on every data pull request.
  • TechEngine runs unit and integration tests, linting and type checks in a separate workflow.
  • Full-data checks cover schema shape, value ranges, slug and reference uniqueness, and required sources.
  • Coverage jobs compare curated records with upstream catalogs and publish gap reports.

Reproduction path

The same dataset can be validated, seeded into a database, served with FastAPI and exported as a static dump. Docker Compose provides PostgreSQL while the default local path uses SQLite.

Known limits

Vendor page changes and region-specific naming still require collector maintenance. No public single-number accuracy score covers the entire dataset, so overall accuracy is marked not measured. Automation verifies provenance and structure but cannot guarantee that an upstream source is itself correct.

05

로드맵

TechAPI의 완료 항목과 이후 확장 방향

완료

  • 다섯 데이터 범주와 공통 경로 규칙
  • 경량 PR 검증과 TechEngine 전체 무결성 검사
  • FastAPI 읽기 API, SQLModel 적재와 정적 JSON 덤프
  • 데이터셋·엔진 사이트와 GitHub Actions 배포
  • 양방향 서브모듈 SHA 동기화

진행 중

  • 외부 카탈로그와 큐레이션 데이터 사이의 누락 항목 보고
  • 주간 benchmark 갱신과 신규 SKU 초안 생성 안정화

계획

  • Intel ARK, AMD 제품 페이지, TechPowerUp 등 검증 가능한 원천 확대
  • 데이터 품질 이력과 알고리즘 버전의 가시성 강화
  • 소비 프로젝트가 변경을 안전하게 감지할 수 있는 버전 정책 정교화

제외

쓰기 권한이 필요한 사용자 계정 API나 상거래 기능은 현재 범위가 아니다. TechAPI는 읽기 가능한 공개 데이터와 검증 파이프라인에 집중한다.

05

Roadmap

Completed capabilities and future direction

Completed

  • Five data categories and common path rules
  • Lightweight pull-request validation and full integrity checks
  • FastAPI read API, SQLModel seeding and static JSON dumps
  • Dataset and engine sites with GitHub Actions delivery
  • Bidirectional submodule revision synchronization

In progress

  • Gap reports between upstream catalogs and curated data
  • Stabilization of scheduled benchmark refresh and new-SKU drafting

Planned

  • Additional verifiable sources such as Intel ARK, AMD product pages and TechPowerUp
  • More visible data-quality history and algorithm versions
  • Clearer version policies for downstream consumers

Out of scope

Authenticated write APIs and commerce workflows are not current goals. TechAPI remains focused on verifiable public data and dependable read delivery.

확대 보기Expanded view

도식을 좌우로 이동하거나 확대해 세부 흐름을 확인할 수 있습니다.

Pan or zoom the diagram to inspect the detailed flow.