SPB Git forge

spb/websensor

Public
33commits 1branches 0releases
3.4 MBsize
maindefault branch
10 days agolast push
TypeScript 55.4% Python 43.2% SQL 1.2%
1.5 KB

# Connector: PDF documents (pdf)

Purpose: official documents that only exist as PDF and whose text matters — framework and standards documents (NIST CSF 2.0, SP 800-53), forms and instructions (SEC Form 10-K, IRS 1040, CRA guides), network rules (Visa Core Rules), licence agreements (Apple SLA), AI-lab policies (OpenAI Preparedness Framework). Only stable URLs qualify; dated file names (statement_20250917.pdf) belong to feed/HTML sensors, not here. Sensor types: FILE, PDF_INDEX. Config: maxPages (40 — extraction cost grows linearly; long documents are sampled from the start), headers, timeoutMs. Extraction: pdf.js through unpdf (pure JS, no native binary); page texts joined, whitespace normalized, then canonicalizeText (same scrubbing as HTML pages) → text diff. Metadata: title, author, producer, creation and modification dates (publishedAt). Confidence: < 40 chars extracted (scanned document) → confidence 0.2 — the pipeline then only reports changed/unchanged on the raw hash, never a text diff. Failure handling: HTML instead of PDF → html_not_pdf; body not starting with %PDF- → bad_pdf. Limit 40 MB. Known quirks: pdf.js prints Warning: TT: undefined function for some embedded fonts (harmless); very large documents (OpenAI framework 27 MB) take several seconds — tier D. Tests: covered live by the validator (config/sources.d/35-documents-data.yaml). Verified: 2026-09-08 on 10 PDFs.