Debug Python Crawler: Restore Data Extraction

Cliente Freelancer · Remoto · Remoto · freelance · mid · 30–250 USD

Publicada el 2026-07-23

Descripción de la oferta

My existing Python crawler ran flawlessly in early tests but now fails to pull the content I need. It should collect three kinds of assets in a single pass—raw HTML, embedded or linked PDF files, and plain-text documents—from a mix of simple static pages and those that render extra material through JavaScript. Pages that sit behind a login are not in scope for now, so authentication flows can be ignored. The problems I see: • Some pages return only partial HTML, losing key sections that appear after JavaScript executes. • PDF links are discovered but not downloaded consistently. • Text files get downloaded, yet their contents arrive empty or garbled. I would like you to: 1. Review the current codebase (requests/BeautifulSoup for static parts, a lightweight Selenium fallback for dynamic ones) and pinpoint where extraction breaks. 2. Patch the logic so it reliably gathers the three asset types above, saving them to the folder structure already defined. 3. Provide a short report or inline comments explaining the fix so I can maintain it later. A successful hand-off means running the updated script against my test URL list and seeing complete, readable HTML, PDFs saved intact, and text files with their full content. If any additional Python libraries are required, please note the exact versions in the report. Let me know if you need sample URLs or current error logs and I’ll share them right away.

Skills

Fuente original: freelancer

Análisis JobHunter