{"repo":"AndreaBozzo/Ceres","free":true,"listed":false,"github":"https://github.com/AndreaBozzo/Ceres","clone":"git clone https://github.com/AndreaBozzo/Ceres.git","description":"Harvest-first Rust toolkit for open data metadata: sync 300+ portals (CKAN, DCAT, SPARQL, data.json, Socrata, OpenDataSoft, ArcGIS, CSW, STAC ) into PostgreSQL, add optional local embeddings and semantic search. Ceres Open Data Index (2M+ datasets) on Hugging Face.","language":"Rust","stars":13,"topics":["metadata","pgvector","rust","harvester","ckan","data-engineering","ollama","cli","data-catalog","dcat"],"license":"Apache-2.0","category":"data-pipelines","readme_excerpt":"Ceres Harvest-first toolkit for open data portals — one synchronized catalog from 300+ portals and 2M+ datasets Website • Quick Start • Supported Portals • The Index • REST API --- Open data is fragmented across thousands of portals speaking different APIs — CKAN, DCAT, SPARQL, data.json , Socrata, OGC CSW, STAC, and SDMX. Ceres harvests them into one synchronized PostgreSQL catalog : incremental sync, delta detection, stale tracking, bounded memory on multi-million-dataset sources. Everything else is layered on top, and optional: local embeddings, semantic search, a REST API, and reproducible Parquet snapshots published as the public Ceres Open Data Index . Named after the Roman goddess of harvest and agriculture. At a Glance - 300+ portals harvested and kept in sync — national portals (data.gov, data.europa.eu, data.slovensko.sk, govdata.de), cities (Milano, NYC, Zurich), and agencies - 2M+ datasets in the live catalog and published snapshots - 10 harvest paths shipped, including OGC CSW 2.0.2, collection-level STAC, and dataflow-level SDMX alongside the major open-data portal families - Metadata-only by default — no embedding provider, API key, or GPU required to build and maintain a catalog - Local-first embeddings via Ollama when you want semantic search; Gemini and OpenAI supported - Reproducible exports — Parquet snapshots with versioned manifests, SHA-256 checksums, coverage/quality reports, and changelogs Why Harvest-First Most open data tooling starts at search. In ","default_branch":null,"files":null,"tree":[],"storefront":"/r/AndreaBozzo","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AndreaBozzo/Ceres/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}