Government Tender LLM Extractor

Cliente Freelancer · Remoto · Remoto · freelance · mid · 750–1250 INR

Publicada el 2026-07-27

Descripción de la oferta

I have a growing archive of government tender documents—PDFs, Word files, scanned images—that need to be automatically broken down into meaningful chunks and then mined for key data points (issuer, deadlines, scope, budget caps, mandatory criteria, contact details, etc.). My end-goal is to wrap this in a “meta-LLM” layer, so the quality of the chunking and field-level extraction must be rock solid and ready to feed downstream reasoning models. Here is what I need from you: • Build a Python-based pipeline that ingests raw tender files, cleans them, splits them into semantically coherent chunks, and extracts the structured information I specify. • The solution has to cope with long documents (100+ pages), mixed layouts, and occasional OCR noise. • Please rely on modern LLM tooling—HuggingFace Transformers, LangChain or LlamaIndex for chunk orchestration, plus a vector store of your choice if it helps internal look-ups. • Deliver clean, well-documented code, a small demo dataset, and a README showing how to run everything locally or in a cloud notebook. To keep the bidding process focused, just point me to past work that proves you have already built or significantly contributed to similar LLM-powered document extraction or RAG pipelines. Screenshots, repos, short videos—whatever best demonstrates results—are welcome. Once the core chunking & extraction is stable, I intend to layer in search and categorisation, so a modular design will earn extra points.

Skills

Fuente original: freelancer