JavaScript 58.4%
Python 26.8%
CSS 7.5%
Objective-C 3%
R 2.6%
HTML 1.7%
1* =============================================================================2* pdb_api.do — utiliser la PDB API (SEC Project Intelligence Database, UQO) dans Stata3*4* Stata ne peut pas envoyer d'en-tête HTTP avec `copy` ou `import delimited`.5* Deux voies :6* A) Stata 16+ avec Python intégré (recommandé) : `python:` appelle l'API et écrit un CSV.7* B) Toute version : `shell curl` télécharge le CSV / JSON, puis `import delimited`.8* Remplacer VOTRE_CLE ci-dessous (ou définir la variable d'environnement PDB_API_KEY).9* =============================================================================10clear all11set more off12global PDB_BASE "https://www.pdb-api.co/v1"13global PDB_KEY "VOTRE_CLE"1415* -----------------------------------------------------------------------------16* A) Voie Python (Stata 16+) — définit trois programmes : pdb_projects, pdb_mentions, pdb_sql17* -----------------------------------------------------------------------------18python:19import csv, json, os, urllib.parse, urllib.request20from sfi import Macro2122BASE = Macro.getGlobal("PDB_BASE"); KEY = os.environ.get("PDB_API_KEY") or Macro.getGlobal("PDB_KEY")2324def _req(path, params=None, body=None):25 url = f"{BASE}/{path}" + (f"?{urllib.parse.urlencode(params)}" if params else "")26 data = json.dumps(body).encode() if body is not None else None27 req = urllib.request.Request(url, data=data, headers={"X-API-Key": KEY, "Content-Type": "application/json"})28 with urllib.request.urlopen(req, timeout=120) as r:29 return json.loads(r.read())3031def _flat(rec):32 return {k: (" | ".join(map(str, v)) if isinstance(v, list) else v) for k, v in rec.items()}3334def _write(rows, path):35 if not rows:36 open(path, "w").close(); return 037 with open(path, "w", newline="", encoding="utf-8") as f:38 w = csv.DictWriter(f, fieldnames=list(rows[0].keys())); w.writeheader()39 for r in rows: w.writerow(_flat(r))40 return len(rows)4142def pdb_pages(path, csvfile, **filters):43 """Tous les enregistrements paginés (500 par page) -> CSV."""44 rows, offset = [], 045 while True:46 page = _req(path, {**filters, "limit": 500, "offset": offset})47 rows += page["items"]; offset += len(page["items"])48 if not page["items"] or offset >= page["total"]: break49 n = _write(rows, csvfile); Macro.setLocal("pdb_n", str(n))5051def pdb_sql(sql, csvfile, limit=5000):52 res = _req("sql", body={"sql": sql, "limit": limit})53 rows = [dict(zip(res["columns"], r)) for r in res["rows"]]54 n = _write(rows, csvfile); Macro.setLocal("pdb_n", str(n))55end5657* --- Programmes Stata enveloppant les fonctions Python ------------------------58capture program drop pdb_projects59program define pdb_projects60 * usage : pdb_projects, filters(type=data_center min_amount=5e8) [file(x.csv)]61 syntax , [FILters(string) FILE(string)]62 if "`file'" == "" local file "pdb_projects.csv"63 local kw ""64 foreach f of local filters {65 gettoken k v : f, parse("=")66 local v = subinstr("`v'", "=", "", 1)67 local kw `"`kw' `k'="`v'","'68 }69 python: pdb_pages("projects", "`file'" `kw')70 import delimited using "`file'", clear varnames(1) encoding(utf8) stringcols(_all)71 destring total_amount_usd n_mentions n_filings first_year last_year avg_confidence, replace force72 gen date_first = date(first_seen, "YMD"); format date_first %td73 gen date_last = date(last_seen, "YMD"); format date_last %td74 di as txt "`pdb_n' projets importés"75end7677capture program drop pdb_mentions78program define pdb_mentions79 syntax , [FILters(string) FILE(string)]80 if "`file'" == "" local file "pdb_mentions.csv"81 local kw ""82 foreach f of local filters {83 gettoken k v : f, parse("=")84 local v = subinstr("`v'", "=", "", 1)85 local kw `"`kw' `k'="`v'","'86 }87 python: pdb_pages("mentions", "`file'" `kw')88 import delimited using "`file'", clear varnames(1) encoding(utf8) stringcols(_all)89 destring amount_usd confidence, replace force90 gen date_filing = date(filing_date, "YMD"); format date_filing %td91 di as txt "`pdb_n' mentions importées"92end9394capture program drop pdb_sql95program define pdb_sql96 * usage : pdb_sql "select ... " [, file(x.csv) limit(5000)]97 syntax anything(everything name=sql) [, FILE(string) LIMit(integer 5000)]98 if "`file'" == "" local file "pdb_sql.csv"99 local sql = subinstr(`"`sql'"', `"""', "", .)100 python: pdb_sql("""`sql'""", "`file'", `limit')101 import delimited using "`file'", clear varnames(1) encoding(utf8)102 di as txt "`pdb_n' lignes importées"103end104105* -----------------------------------------------------------------------------106* Exemples (voie A)107* -----------------------------------------------------------------------------108109* 1. Centres de données > 500 M$110pdb_projects, filters(type=data_center min_amount=5e8 sort=amount)111list ticker project_name canonical_location total_amount_usd first_seen in 1/10, clean112113* 2. Tout un secteur, puis agrégation114pdb_projects, filters(sector=Utilities) file(utilities.csv)115tab project_type, sort116collapse (count) n=project_id (sum) capital=total_amount_usd, by(project_type)117gsort -n118list, clean119120* 3. Mentions 8-K sur l'IA depuis 2024121pdb_mentions, filters(type=ai_initiative form=8-K date_from=2024-01-01)122gen yr = year(date_filing)123tab yr124125* 4. SQL libre : mentions IA par année126pdb_sql "select year(filing_date) as yr, count(*) n from project_mentions where project_type = 'ai_initiative' group by 1 order by 1"127twoway connected n yr, title("Initiatives IA dans les filings") ytitle("mentions")128129* 5. Panel entreprise x année pour l'économétrie130pdb_sql "select cik, any_value(ticker) ticker, any_value(sector) sector, year(first_seen) as yr, count(*) n_projets, sum(total_amount_usd) capital_usd from projects group by 1, 4 order by 1, 4", file(panel.csv)131encode cik, gen(id)132xtset id yr133xtpoisson n_projets i.yr, fe // intensité d'annonce de projets, effets fixes entreprise134* xtreg ln_capital i.yr, fe après : gen ln_capital = ln(capital_usd)135136* -----------------------------------------------------------------------------137* B) Voie curl (toute version de Stata) : CSV d'export puis import delimited138* -----------------------------------------------------------------------------139* Sur macOS / Linux : `shell` ; sur Windows : remplacer par `winexec` ou `shell curl.exe ...`140shell curl -s "$PDB_BASE/projects/export.csv?type=plant_construction" -H "X-API-Key: $PDB_KEY" -o usines.csv141import delimited using "usines.csv", clear varnames(1) encoding(utf8) stringcols(_all)142destring total_amount_usd n_mentions n_filings avg_confidence, replace force143gen date_first = date(first_seen, "YMD"); format date_first %td144describe, short145summarize total_amount_usd, detail146147* SQL via curl : la réponse est en JSON ; convertir en CSV avec python3 du système148shell curl -s -X POST "$PDB_BASE/sql" -H "X-API-Key: $PDB_KEY" -H "Content-Type: application/json" -d "{\"sql\":\"select project_type, count(*) n, sum(total_amount_usd) capital from projects group by 1 order by n desc\",\"limit\":100}" | python3 -c "import csv,json,sys; d=json.load(sys.stdin); w=csv.writer(sys.stdout); w.writerow(d['columns']); w.writerows(d['rows'])" > types.csv149import delimited using "types.csv", clear varnames(1)150list, clean151