Urgent News Site Data Scrape
Publicada el 2026-07-15
Descripción de la oferta
I need an automated scraper that gathers data from several news sites in near real-time. The tool should loop through a list of URLs I will provide, respect each site’s robots.txt where possible, and export the captured information to CSV or JSON so I can feed it straight into my analysis pipeline. I’ll share the exact fields during kickoff, but the scraper must be flexible enough to handle common article elements—headline, body text, author byline, publication date, and source URL—and easy to extend if I add more outlets later. Time is critical. Delivery within 24–48 hours is preferred, so please lean on a proven stack such as Python with Scrapy/BeautifulSoup, Node with Cheerio, or any robust alternative you already master. The script should: • Rotate user agents and accept a proxy list to avoid blocks • Log failed requests for easy reruns • Be clearly commented and organized so I can update selectors myself Deliverables 1. Executable script or notebook with all dependencies noted 2. Sample output file containing at least 50 successfully scraped articles 3. Short README explaining setup, configuration, and how to add new sites The project is complete once I can run the scraper locally and reproduce your sample output without errors.
Skills
Fuente original: freelancer