rate-foundry

Bulk export

For analysis over the whole catalog, use the weekly parquet snapshot instead of paging the API.

A scheduled job reads the catalog through the public API and commits a snapshot to the repository: exports/rates-YYYY-MM-DD.parquet with a matching exports/manifest-YYYY-MM-DD.json (record count, API base, parquet SHA-256). Columns match the API record fields; provenance and path_energies_eV are JSON strings.

import httpx

# The same data over the API, paginated 1000 at a time:
BASE = "https://rate-foundry-api.dirac-e08.workers.dev"
rows, cursor = [], None
while True:
    params = {"limit": 1000, **({"cursor": cursor} if cursor else {})}
    page = httpx.get(f"{BASE}/v1/rates", params=params, headers={"User-Agent": "curl/8"}).json()
    rows += page["data"]
    cursor = page["page"]["next_cursor"]
    if not cursor:
        break
print(len(rows), "records")
assert rows

Verify a snapshot by comparing the SHA-256 of the parquet file to the manifest.