1. Collection and mention contract
Naver news and blog, RSS and optional Tavily collectors normalize different documents into mentions with title, body, source, author, publication time and URL. Source cursors and duplicate keys reduce repeated ingestion, while a failed source does not discard successful collectors. Original evidence and collection time remain available to generated reports.
2. Windows and feature generation
The feature pipeline computes volume, growth, sentiment, keyword, user and propagation groups per brand window. Recent input and a 29-day baseline remain distinct, with explicit missing-value and minimum-sample behavior. Feature vectors and baseline versions are stored so a decision can be reconstructed.
3. Combined detection and severity gate
Isolation Forest emits a multivariate anomaly score, while domain rules evaluate explainable conditions such as growth, negative ratio and risk keywords. A combiner maps both outputs to normal, caution or alert. Model-only and rule-only triggers remain distinct evidence in the report.
4. LangGraph and human review
Normal detections are recorded and stop. Caution and alert enter LangGraph nodes for evidence selection, cause hypotheses, response options and report assembly. Nothing reaches an external channel before Human Review approves, revises or rejects the report. Approved reports go to Slack or email, with a console sink when channels are absent.
5. Storage, access and operations
A repository adapter absorbs SQLite and PostgreSQL placeholder, auto-increment and UPSERT differences. CLI, Streamlit, scheduler and FastAPI /analyze, /history, /reports and /monitor use the same application service. Collector, LLM and notifier failures become stage status while persisted detections and review records remain intact.
6. State and data-model ownership
Shared mention, window and detection models form the collector-detector contract, while RiskState owns state during LangGraph execution. The repository adapter owns persistence so API, dashboard and scheduler remain unaware of SQL dialect differences.
Collection fan-out, sentiment processing, Isolation Forest fitting and LLM latency are primary bottlenecks. Per-channel timeouts and partial success, model reuse, bounded windows and idempotent incident writes require validation; a rule-based report remains the LLM fallback.
8. Security, observability and debt
API keys are isolated in environment variables, while collector input, notification targets and dashboard access require validation. Stage status and persisted records provide audit evidence, but distributed tracing, SLOs and production-send approval are missing or planned. Production precision and recall have not been measured.