# Company Atlas — Product Specification (master brief, verbatim from the founder, 2026-09-12) ## 0. PROJECT IDENTITY **Product name:** Company Atlas **Primary domain:** `https://www.company-atlas.com` **Canonical hostname:** `www.company-atlas.com` **Product category:** Global company intelligence / continuous company monitoring / historical corporate data platform **Core concept:** Build the continuously updated historical record of the world's companies. Company Atlas is NOT a static company directory. Company Atlas is a global intelligence platform composed of **thousands, and eventually millions, of persistent public-web connectors/sensors** attached to companies around the world. Each connector continuously observes a specific public surface of a company: corporate website, careers pages, job boards, newsroom, investor relations, pricing, products, services, documentation, APIs, changelogs, leadership pages, office/location pages, sustainability pages, partnerships, customer stories, legal pages, terms, privacy policies, status pages, support portals, blogs, research pages, GitHub/public development activity where appropriate, public structured feeds, sitemaps, public subdomains, other legally/publicly accessible corporate surfaces. The platform must continuously turn these observations into: 1. raw observations, 2. normalized snapshots, 3. structural fingerprints, 4. detected changes, 5. structured corporate events, 6. historical timelines, 7. proprietary metrics, 8. comparative intelligence, 9. search, 10. APIs, 11. alerts, 12. datasets. The core competitive advantage is **time**. Every day the platform operates, its historical dataset becomes harder to reproduce. > Anyone can crawl a website today. Very few can reconstruct exactly how 100,000 companies changed over the previous five years. Company Atlas should become that historical record. # 1. NORTH STAR > **Build the most comprehensive continuously updated machine-readable history of how companies evolve.** We are not merely collecting pages. We are measuring corporate change. Examples: Which companies are accelerating hiring? Cutting hiring? Changed pricing? Launched products? Quietly removed products? Entered new countries? Added AI-related roles? Changed executive leadership? Expanding developer ecosystems? Changing positioning? Moving toward enterprise customers? Changed their terms? Opened or closed offices? Increasing public communication activity? Which industries are changing fastest? Which geographies attract expansion? Which companies appear to be preparing launches? Which businesses suddenly have abnormal website activity? # 2. PRODUCT PRINCIPLES ## 2.1 Historical-first Never overwrite history when a new version arrives. Changes create new versions. We care about state(t0), state(t1), state(t2)… not just current_state. Current state is simply the latest historical state. ## 2.2 Everything important should be timestamped discovered_at, first_seen_at, last_seen_at, fetched_at, changed_at, published_at, effective_at, processed_at — depending on entity type. ## 2.3 Raw data and interpreted data must remain separate Never destroy source observations after interpretation. RAW SOURCE → RAW SNAPSHOT → NORMALIZED SNAPSHOT → DIFF → EVENT → LLM ENRICHMENT → METRICS. This allows future reprocessing when extraction models improve. ## 2.4 LLMs are enrichment engines, not the crawler Do NOT send every fetched webpage to an expensive model. fetch → normalize → hash → compare → structural diff → semantic-change candidate → LLM only if useful. ## 2.5 Build reusable connector families Do NOT hard-code every company individually. The system combines generic connectors + site adapters + discovered endpoints + company-specific configuration and must be capable of creating thousands of connectors automatically. ## 2.6 Every crawl should answer a question Did pricing change? Did the leadership team change? Did jobs increase? Was a product introduced? Were locations added? Did documentation change? Specialized sensors produce better data than storing full websites. # 3. SCALE TARGETS Phase 1: 5,000 companies / 25,000+ active sensors. Phase 2: 25,000 / 150,000+. Phase 3: 100,000 / 750,000+. Phase 4: 1,000,000 / 5M–15M. An important company could have 10–100 monitored surfaces; long-tail companies 2–10. Sensors are lightweight entities. # 4. TERMINOLOGY **Company** canonical organization (Stripe, Inc.). **Domain** canonical corporate domain (stripe.com). **Surface** logical information surface (careers, pricing, newsroom, investor_relations, products, leadership, locations). **Connector** a software strategy able to collect a type of source (generic_html_connector, sitemap_connector, greenhouse_connector, lever_connector, json_endpoint_connector, rss_connector). **Sensor** a deployed connector instance attached to a specific source (Company: Stripe · Surface: Careers · Connector: generic_careers · URL: https://stripe.com/jobs) — sensors execute repeatedly. **Observation** one fetch/read of a sensor. **Snapshot** normalized representation of an observation. **Change** difference between snapshots. **Event** a meaningful interpreted corporate change. # 5. HIGH-LEVEL ARCHITECTURE Company Registry / Graph → Discovery Engine + Scheduler → Connector Workers (HTTP, Browser, Feeds) → Raw Storage → Normalization → Fingerprint / Diff → Change Detector → (insignificant → archive | meaningful → Event Pipeline → LLM Enrichment → Event Store → Metrics Engine → API / Search / UI). # 6. COMPANY REGISTRY Canonical company entity: id, slug, legal_name, display_name, canonical_domain, website, description, industry[], country, headquarters, founded_year, company_type, public_company, ticker, status, created_at, updated_at. Keep factual provenance. Do not merge entities based only on similar names. # 7. COMPANY GRAPH Relationships: company → domain, brand, parent, subsidiary, executive, office, product, competitor, acquisition, investor, technology, job, industry, country, announcement, event. Every edge supports temporal validity where relevant (valid_from, valid_to, first_seen, last_seen, source). # 8. CONNECTOR ARCHITECTURE Thousands of connectors but not thousands of unrelated codebases. Base interface: discover / fetch / extract / normalize / fingerprint / compare / emit_events. Metadata: connector_id, name, version, category, fetch_mode, supports_discovery, supports_incremental. # 9. CONNECTOR FAMILIES 9.1 Core website connectors: homepage, about, company, leadership, team, products, services, solutions, industries, customers, partners, locations, contact, pricing, documentation, developer, API, changelog, blog, news, press, research, careers, legal, terms, privacy, security, trust, sustainability, ESG, investor relations. 9.2 Discovery connectors: robots.txt, sitemap.xml, sitemap index, RSS, Atom, navigation discovery, HTML link graph, JSON-LD, schema.org, alternate language URLs, subdomain discovery, known URL patterns. 9.3 Job connectors: Greenhouse-like, Lever-like, Workday public career surfaces, SmartRecruiters-like, Ashby-like, custom career pages, JSON-backed boards, HTML listings. Avoid relying on a single vendor API; prefer public endpoints already used by public pages when lawful and appropriate. # 10–11. SENSOR CREATION & AUTOMATIC DISCOVERY canonical domain → homepage fetch → robots inspection → sitemap discovery → navigation analysis → subdomain discovery → URL classification → surface identification → sensor creation. Signals: anchor text, URL paths, navigation hierarchy, page titles, schema markup, sitemap metadata, headings, CMS conventions, subdomains. Classes: CAREERS, NEWSROOM, BLOG, PRODUCTS, PRICING, ABOUT, LEADERSHIP, LOCATIONS, INVESTOR_RELATIONS, DOCUMENTATION, CHANGELOG, LEGAL, OTHER — each with a confidence. # 12. SENSOR QUALITY SCORE source reliability × extraction confidence × historical stability × semantic importance × recency → quality_score 0–100. Low quality sensors may be automatically reviewed. # 13. FETCH MODES A — lightweight HTTP (default: HTML, JSON, XML, RSS, sitemaps). B — browser rendering only when needed, pooled. C — public frontend network endpoint extraction where a public page loads data from public endpoints; store provenance. Never bypass authentication, access controls, CAPTCHAs; never collect non-public information. # 14. POLITENESS / COMPLIANCE Domain-specific rate limiting, robots policy awareness, retry budgets, crawl-delay, backoff, concurrency caps, clear user agent, attribution, provenance, suppression mechanism, legal/compliance flags. Never collect content requiring unauthorized access; never circumvent authentication; no private customer data. # 15–16. SCHEDULER & ADAPTIVE CRAWLING Tiers: A 5–15 min · B 30–60 min · C 6 h · D 24 h · E 3–7 days. Frequency adapts: significant changes → temporary burst (1/15 min) then decay. next_interval = base × stability × importance × failure × activity. # 17–20. CHANGE DETECTION, NORMALIZATION, BLOCK DIFFING, SIGNIFICANCE Noise (dates, analytics IDs, rotating testimonials, randomized content, tracking params, ads, cookie banners, session ids) is normalized away. Hashes: raw_hash, normalized_hash, structural_hash, semantic_hash. Normalization: remove scripts/styles, normalize whitespace, remove dynamic attributes, normalize URLs, strip tracking/session params, canonicalize headings, extract primary content, preserve structured data → normalized DOM, plaintext, semantic blocks, structured fields. Block-level diffing over semantic blocks (header, hero, product card, pricing plan, job listing, executive bio, office location, news article, FAQ, table) with stable identities → added / removed / modified / moved. Significance: 0–0.20 noise · 0.20–0.40 minor · 0.40–0.65 meaningful · 0.65–0.85 major · 0.85–1.00 critical. Inputs: % text changed, semantic similarity, page importance, affected structured entities, novelty, cross-source confirmation, historical baseline. # 21–23. EVENT TAXONOMY, MODEL, DEDUPLICATION Types: PRODUCT, PRICING, HIRING, LEADERSHIP, LOCATION, FINANCING, M&A, PARTNERSHIP, STRATEGY, TECHNOLOGY, LEGAL, MARKETING, DEVELOPER, SECURITY, OPERATIONS, SUSTAINABILITY, INVESTOR_RELATIONS, COMMUNICATION, OTHER. Subtypes: PRODUCT_LAUNCH, PRODUCT_REMOVAL, PRODUCT_RENAME, PRICE_INCREASE, PRICE_DECREASE, NEW_PRICING_TIER, JOB_COUNT_INCREASE, JOB_COUNT_DECREASE, NEW_EXECUTIVE, EXECUTIVE_REMOVED, NEW_OFFICE, OFFICE_REMOVED, COUNTRY_EXPANSION, NEW_PARTNERSHIP, ACQUISITION, DIVESTITURE, API_LAUNCH, DOCUMENTATION_CHANGE, TERMS_CHANGE, BRAND_REPOSITIONING. Event model: id, company_id, event_type, event_subtype, importance, confidence, title, summary, old_value, new_value, detected_at, effective_at, source_url, sensor_id, snapshot_before, snapshot_after, model_version. The same corporate event appearing on homepage, press release, blog, pricing page, IR → clustered into one canonical event with sources; corroboration increases confidence. # 24–25. LLM ENRICHMENT & MODEL ABSTRACTION LLMs receive only relevant changed content (company metadata, source type, before blocks, after blocks, structured changes). Tasks: classification, summary, importance, entity extraction, structured event generation, industry tagging, sentiment where appropriate, strategy interpretation. Structured JSON validated against schemas. Never hard-code one vendor: LLMProvider / EmbeddingProvider / RerankerProvider; local models, OpenAI-compatible endpoints; work with local cluster inference wherever practical. # 26–28. STORAGE Every relevant version retained (sensor_id, url, fetched_at, status_code, content_hash, normalized_hash, storage_pointer, content_type, size_bytes); compressed object storage; dedupe by hash. Layers: PostgreSQL (entities/metadata), object storage (raw), columnar analytics, search engine, vector index — abstracted. Internal event bus topics: company.created, sensor.created, sensor.fetch.requested/completed, snapshot.created, change.detected, event.generated, metric.updated, alert.triggered; async idempotent workers. # 29–37. PROPRIETARY METRICS Company Activity Score (0–100: website changes, product changes, news frequency, job movement, leadership changes, documentation, pricing). Hiring Momentum (open jobs, new/removed, departments, locations, seniority, remote ratio, skills; 7/30/90-day, YoY). AI Adoption Score (observable public signals only: AI products, jobs, docs, marketing, partnerships, research, leadership roles — never claim internal use without evidence). Product Velocity. Geographic Expansion Score. Developer Momentum. Corporate Change Index = 0.25 hiring + 0.20 product + 0.15 geographic + 0.15 leadership + 0.10 developer + 0.10 communication + 0.05 pricing (weights to become empirical). Unusual Activity Detection from per-company baselines ("Companies behaving unusually today"). Cross-company signals (AI hiring acceleration by industry, SaaS pricing increases, US manufacturing expansion…). # 38–48. PRODUCT SURFACES Industry Atlas (living index per industry: activity, companies, events, hiring, expansion, trending). Country Atlas (/country/canada …: activity, hiring, expansion, industry mix, top movers, new entrants). Company profile `/company/stripe`: header with Activity Score, Hiring Momentum, Product Velocity, AI Adoption; sections Overview, Timeline, Signals, Jobs, Products, Locations, Leadership, Technology, Sources, Historical. Company Timeline (dated entries, filters all/products/jobs/pricing/leadership/locations/legal/news/developer). Historical Page Viewer (semantic diff between versions). Global Live Feed on the homepage (cards: company, change, "2 minutes ago", incremental updates). Homepage hero **The Live Atlas of Global Companies** with animated counters (companies, sensors, observations, changes, structured events) and sections: Hero, Live Activity Feed, Companies Moving Fastest, Global Activity Map, Hiring Momentum, Product Launches, Pricing Changes, AI Adoption, Industries, Countries, Trending Signals, Platform Statistics, API/Data CTA. Map with clustering (HQs, expansions, offices, event density, hiring). Search over companies, industries, events, products, people, locations, keywords, technologies ("companies hiring AI engineers in Canada"). Later: "Ask Company Atlas" natural-language search always linking back to events and sources. # 49–50. PROVENANCE & CONFIDENCE Every important fact traceable (Source, First seen, Last checked, Historical evidence, Confidence). Statuses: VERIFIED, HIGH CONFIDENCE, LIKELY, INFERRED, LOW CONFIDENCE. Never present inference as verified fact without labeling. # 51–54. API, EVENT API, STREAMING, EXPORTS API-first `/api/v1`: companies, companies/{id}, events, metrics, jobs, history, events (filters event_type, country, since…), industries, countries, search. Streaming later (WebSocket, SSE, webhooks). Exports JSON, CSV, Parquet, NDJSON. # 55–56. USERS & WATCHLISTS Public browsing without account. Accounts only for watchlists, alerts, API access, saved searches, dashboards, exports. Watchlists with alerts (product launch, pricing change, leadership change, hiring spike, location expansion, activity anomaly). # 57–66. INTERNAL ADMIN & OPERATIONS /admin modules: companies, domains, sensors, connectors, crawl queue, worker health, failures, snapshots, changes, events, duplicates, LLM jobs, metrics, sources, alerts, system statistics. Connector Control Center (version, active sensors, success rate, latency, change detection rate, errors, last deployment). Sensor Control Center (filters healthy/failing/stale/blocked/redirected/low quality/high activity; actions pause/resume/retry/rediscover/change frequency/change connector). Auto-repair (retry → re-fetch sitemap → rediscover navigation → replacement URL → compare content identity → migrate sensor; flag if uncertain). Connector versioning (generic-pricing-v1 → v2; version stored with observations). Failure classes DNS, TIMEOUT, HTTP_4XX, HTTP_5XX, BOT_CHALLENGE, PARSING, SCHEMA, REDIRECT, PAGE_REMOVED, RATE_LIMIT, UNKNOWN with distinct retry policies. Domain budgets (max_concurrency, requests_per_minute, daily_budget, browser_budget, retry_budget). Priority queue (importance, watchlists, recent activity, time since last crawl, sensor value, reliability, cost). Cost accounting per company/sensor/connector/fetch/browser/LLM/GB/event. Observability (fetch/sec, success rate, changes/sec, meaningful event rate, queue lag, browser utilization, storage growth, LLM jobs, false positive rate, connector failures). # 67–69. INFRASTRUCTURE Logical services: atlas-web, atlas-api, atlas-scheduler, atlas-discovery, atlas-fetch-http, atlas-fetch-browser, atlas-normalizer, atlas-diff, atlas-events, atlas-llm, atlas-metrics, atlas-search, atlas-admin, atlas-worker-manager. Stateless workers, horizontal scaling (add machine → register worker → joins queue), configurable node roles. # 70–82. DATA MODEL DETAILS Tables: companies, domains, company_aliases, company_relationships, surfaces, connectors, sensors, observations, snapshots, changes, events, event_sources, people, jobs, products, locations, metrics, metric_series, crawl_runs, failures, users, watchlists, alerts. Metrics are time series (never only latest). Jobs: title, department, location, remote, employment_type, seniority, skills, salary where public, posted_at, first_seen, last_seen, removed_at. Leadership: name, title, role category, first/last seen, source — disappearance is "No longer listed on monitored leadership page", never "Fired". Products: new/changed/renamed/removed from public catalog. Pricing: plan, currency, billing period, price, features, usage units, enterprise/contact sales — preserve every version. Locations: office/store/factory/warehouse/lab/headquarters with city/region/country — never invent exact addresses. Technology signals: observed / stated / inferred. Newsroom: title, URL, published_at, category, entities, summary (first-party = high provenance). Investor relations for public companies (not a substitute for regulated filings). Legal change tracking with semantic summaries ("Section 7 was materially updated"), source versions preserved. Duplicate company resolution by multiple signals; never auto-merge on name alone. Internationalization: store original language + normalized English summaries (source_language, normalized_language, translation_model). # 83–88. URL DESIGN, SEO, DESIGN LANGUAGE Clean URLs: /company/apple, /industry/artificial-intelligence, /country/canada, /events, /live, /rankings, /search, /api, /about. SEO: metadata, OpenGraph, canonical, schema markup, sitemaps; do not index thin profiles until sufficient data. Design: premium intelligence terminal blended with a modern data atlas — dense but readable, beautiful typography, live counters, micro visualizations, timelines, maps, tables, sparklines, confidence indicators, real-time status; avoid generic startup cards, excessive gradients, cartoon visuals, unused whitespace, clutter. Desktop: powerful tables, keyboard search, filters, multi-column, comparison, density. Mobile first-class: responsive typography, bottom navigation (Home, Live, Search, Rankings, Watchlist), touch controls, sticky search, compact metrics, collapsible filters, swipe-friendly timelines, fast loads. Live visual language: green live dot, "17 sec ago", new-event animation, counter increments, sparkline updates — never chaotic. # 89–93. RANKINGS, COMPARISON, GLOBAL INDEX, SIGNALS, PREDICTIONS Rankings: Most Active, Fastest Hiring Growth/Decline, Highest Product Velocity, Most AI-Active, Fastest Geographic Expansion, Highest Developer Momentum, Most Pricing Changes, Most Unusual Activity — windows 24h/7d/30d/90d/1y. Comparison `/company/compare?companies=stripe,adyen,block` (activity, hiring, product velocity, AI, locations, events). Global Corporate Activity Index (baseline 100 = normalized historical activity; by industry, country, size, event type). Signals (hiring surge, hiring freeze, launch buildup, international expansion, pricing migration, developer push, enterprise repositioning, AI acceleration) labeled as signals, not facts. Predictions only as explicitly probabilistic ("Possible launch preparation signal — Confidence 63%"). # 94–98. FLYWHEEL, MOAT, RETENTION, CAS, BACKFILL More companies → sensors → observations → baselines → anomalies → events → metrics → users → watchlists → prioritization → dataset value. Moat = historical snapshots, normalized entities, change events, sensor reliability history, historical metrics, comparisons, graph. Long-term retention; identical snapshots share content objects; keep observation metadata. Content-addressable storage sha256 `objects/ab/cd/hash`. Backfills marked collection_method = backfill, never blurred with live. # 99–106. SEED, FIRST TARGET, CONNECTOR FACTORY, ONBOARDING, REGISTRY, FIXTURES, ADAPTERS, SDK Seed diversified companies (technology, AI, finance, banking, insurance, retail, energy, manufacturing, healthcare, biotech, pharma, transportation, aerospace, telecom, media, real estate, construction, logistics, automotive, consumer) across US, Canada, Europe, UK, Japan, South Korea, India, Australia, Latin America, Middle East, Africa, Southeast Asia. First target 5,000+ companies / 25,000+ functioning sensors. Bootstrap connector factory: domain → surfaces[{type, url, connector, confidence, frequency}]. Mass onboarding pipeline fully parallelizable (import → canonicalization → domain validation → discovery → surface mapping → sensor generation → initial fetch → quality validation → scheduler activation). Machine-readable connector registry (manifest, schema, implementation, fixtures, tests, version). Fixtures for every major connector; never depend entirely on live sites in CI. Site-specific adapters inherit generic logic. Connector SDK decorator style. # 107–112. EXTRACTION & BANDWIDTH Semantic extraction (DOM, ARIA, headings, lists, tables, JSON-LD, microdata, embedded JSON) over brittle selectors. Structured data first. Crawl fingerprinting (etag, last-modified, content-length, status, hash, DOM fingerprint) with conditional requests. Compression, connection pooling, streaming, size limits, selective HEAD. Do not download images/videos/fonts/binaries/tracking scripts by default. Screenshots only for high-impact changes/human review. # 113–131. IMPORTANCE, SECURITY, CANONICALIZATION, LOOPS, PAGE VALUE, QUALITY, REVIEW, AUDIT, VERSIONING, BACKPRESSURE, IDEMPOTENCY, LOCKS, HEALTH, PUBLIC STATS Event importance ≠ confidence; company importance affects crawl priority only. Security: strict URL validation, SSRF protection, DNS rebinding protection, private-IP blocking, download/redirect limits, content-type checks, sandboxed parsing, secret separation; crawler must NEVER access localhost, RFC1918, metadata endpoints, cluster internal domains, Tailscale/private hosts, file://, unix sockets. URL canonicalization (scheme, host casing, ports, tracking params, fragments, duplicate slashes, session params) with the original preserved. Loop prevention (calendars, facets, random params, endless pagination, session IDs) via budgets. Page value classifier for discovered URLs. Data quality dashboard (coverage, freshness, accuracy, duplicate rate, event confidence, stale/failed sensors, unknown surfaces). Human review queues (company merge, major event, low-confidence extraction, sensor migration, legal-sensitive inference, unexpected activity) feeding evaluation data. Auditability: what changed, where, when, evidence, extraction version, model. Model + prompt versioning (prompts in code, e.g. prompts/event-classifier/v3.md). Reprocessing without recollecting. Backpressure prioritizes important/high-value/fresh/active work without losing queued work. Job idempotency (sensor_id + scheduled_window). Distributed locks. /health /ready /metrics per service. Public /system page with aggregate health only. Public statistics counters. # 132–148. VISUALIZATIONS, THEMES, PERFORMANCE, CACHING, RATE LIMITS, AUTH, PRIVACY, ALERTS, DIGESTS, TRENDS, COVERAGE NORMALIZATION, BASELINES, CALIBRATION, REGRESSION Sparklines, time series, heatmaps, maps, histograms, rankings, timelines, network graphs (few pie charts). Dark/light both intentional (default follows system). Fast SSR, progressive hydration, cached public pages, optimized queries. Cache layers (request, query, entity, search, homepage aggregate) — not live feeds. API rate-limit tiers anonymous/authenticated/paid/internal. Auth (if added): passkeys, OAuth, magic links. Privacy: no consumer profiling; executive data limited to legitimate public professional context. Alerts (activity > 80, new executive, price change, job count −30 %, new country, product launch) via email/web/webhook. Weekly company digest and market digests. Trend engine with rolling windows and momentum (7/30/90 d) normalized by coverage growth. **Coverage normalization is critical**: long-term indices adjust for sensors, company population, crawl frequency, industry/country coverage. Baseline engine per sensor/company (changes/week, jobs, announcements, volatility). Quality calibration (correct / duplicate / noise / misclassified). Regression tests against historical fixtures before changing normalizers, diff engine, extractors, prompts. # 149–155. STACK, REPO, ENV, DOMAIN, DEPLOYMENT, BACKUPS Next.js/TypeScript/React front; Python and/or TypeScript services; PostgreSQL; Redis-backed queue initially or durable broker; S3-compatible object storage; columnar analytics as scale requires; full-text search; Dockerized services; no vendor coupling. Monorepo apps/ services/ packages/ connectors/ prompts/ fixtures/ scripts/ infra/ docs/. `.env.example`, never commit secrets (DATABASE_URL, REDIS_URL, OBJECT_STORAGE_*, LLM_ENDPOINT, LLM_API_KEY, SEARCH_URL). Production hostname www.company-atlas.com with apex redirect, HTTPS, no hard-coded IPs. Rolling updates, health checks, restart policies, persistent storage separation, logs, metrics, service discovery, environment separation. Back up PostgreSQL, configuration, indexes, connector metadata; object storage redundancy; test restores; object storage is not a backup — a second independent copy is required. # 156–161. MVP & PHASES MVP: 5,000 companies, 25,000 sensors, continuous monitoring, company profiles, live feed, historical timeline, change detection, event extraction, search, rankings, activity score, admin monitoring. MVP connector categories: homepage, careers, news, products, pricing, leadership, locations, blog, documentation, investor relations, sitemap. MVP events: NEW_JOB, JOB_REMOVED, JOB_COUNT_CHANGE, NEW_PRODUCT, PRODUCT_REMOVED, PRICING_CHANGE, LEADERSHIP_CHANGE, NEW_LOCATION, NEWS_RELEASE, DOC_CHANGE. Phase 2: company graph, comparisons, industry/country indices, watchlists, alerts, API, historical viewer, AI adoption, developer momentum. Phase 3: 100k+ companies, advanced graph, anomaly engine, NL search, cross-company trends, datasets, streaming API, historical analytics. Phase 4: 1M+ companies, 5M+ sensors. # 162–168. COPY, POSITIONING, BRAND, ERROR UX, TRANSPARENCY, ETHICS, LANGUAGE Headline **The Live Atlas of Global Companies**. Supporting: "Company Atlas continuously observes the public web to track how companies evolve — products, hiring, pricing, leadership, locations, technology, strategy and more." Alt: "Thousands of companies. Millions of observations. One continuously growing historical record." Positioning: not a directory, not a news aggregator, not a web archive, not a financial terminal — **a continuously updated corporate observation network**. Brand: intelligent, global, precise, technical, credible, alive, data-rich, premium; not sci-fi. Error UX: "No monitored evidence available yet." / "Last successfully checked 3 days ago." — never fabricate. Every user-facing event links to its public source; if the source disappears, say so. Ethical inference: "72 monitored listings are no longer visible", never "Company laid off 72 employees". Language: detected, observed, publicly listed, no longer listed, appears, signal, inferred, confirmed by. # 169–181. STATES, PROVENANCE GRAPH, SOURCE UI, TIME ZONES, IDS, SOFT DELETES, CORRECTIONS, SCHEMA EVOLUTION, INDEX REPRODUCIBILITY, NO MAGIC NUMBERS, FEATURE FLAGS, EXPERIMENTATION, DOCS Company states ACTIVE, POSSIBLY_INACTIVE, WEBSITE_UNAVAILABLE, ACQUIRED, DISSOLVED, UNKNOWN (require corroboration). Provenance graph entity field → extraction → snapshot → observation → sensor → URL. Event source UI lists each source with detection times. UTC storage, local rendering, preserve source timezone. Immutable internal ids; slugs may change; never name as key. Soft deletes via statuses. Corrections: status = retracted with audit history. Versioned schemas and careful migrations. Reproducible indices (formula version, inputs, normalization window, computed_at). No magic numbers (typed config). Feature flags for experimental systems. A/B evaluation of extractors/normalizers/classifiers/schedules (unstable metrics labeled). Docs: architecture, connector-sdk, event-taxonomy, data-model, deployment, operations, scoring. # 182–183. CLAUDE CODE OPERATING RULES & STYLE 1. Inspect existing architecture first. 2. Preserve working functionality. 3. Avoid unnecessary rewrites. 4. Prefer reusable abstractions. 5. Add tests with new connector behavior. 6. Keep schemas explicit. 7. Never introduce fabricated data. 8. Never hard-code secrets. 9. Do not break mobile. 10. Do not break historical reproducibility. 11. Do not delete historical data during migrations without explicit reason. 12. Run lint/typecheck/tests after meaningful changes. Favor small composable services, typed contracts, schemas, idempotent workers, deterministic processing, clear observability; avoid giant functions, silent error handling, unbounded crawling, LLMs for deterministic tasks, vendor lock-in. # 184–200. PHILOSOPHY & FINAL DEFINITION Every company on Earth should feel like it has virtual sensors attached (APPLE 63 sensors, NVIDIA 48, STRIPE 41, LOCAL BUSINESS 5) living continuously. "Thousands of connectors" is not marketing: 25,000 deployed sensors initially, then 100,000, 1,000,000, 10,000,000+ — configuration-driven and auto-discoverable. Track dataset age and oldest continuous history prominently. Historical Completeness Score per company (sensor uptime, continuity, source coverage, failed periods). Show data density (active sensors, observations, historical changes, structured events). Discovery feedback loop (event → URL discovery → new sensor). Sensor retirement keeps history and seeks successors. Domain migration continuity. Acquisition continuity (ACQUIRED_BY relationship, no blind merges). Entity resolution with stored confidence. Published API docs (/api) with examples, schemas, pagination, rate limits, filters, webhooks. Monetization readiness (free public, pro research, enterprise API, bulk datasets, alerts, custom monitoring) without crippling the free product. Data licensing metadata; redistribute derived data/changes/facts/metadata rather than raw content where required. Long-term: 1M+ companies, 10M+ sensors, billions of observations, years of continuous history. > **Company Atlas is a distributed global sensor network for companies.** Each company receives persistent monitoring sensors. Each sensor produces observations. Observations become historical snapshots. Snapshots produce changes. Changes become structured events. Events become metrics. Metrics become intelligence. And every day the platform runs, the dataset becomes more valuable. Optimize for what this dataset becomes after 1, 3, 5, 10 years. That accumulated history is the product.