A provenance-preserving corpus of Lok Sabha and Rajya Sabha documents for non-commercial research. Publications are uploaded in bounded tranches and do not imply completeness unless a release explicitly says so. The sources are the Digital Sansad question APIs and the Parliament eLibrary. Each document is mapped to an official source record and its original PDF bytes by SHA-256. The archive retains additional eLibrary ORIGINAL-bundle PDFs, including Hindi variants when offered; those attachments are linked in the manifest but are not yet separately text-extracted. Each tranche provides selected official PDFs, page-level Markdown, plain text and JSON in WebDataset TAR shards…
Independent publisher
Bhuvan
bebhuvan1
Models in Library0
Datasets in Library1
Models on Hugging Face—
Followers—