Complete Lottery Prediction and Betting Automation System (Spain and Brazil) – Continuous Learning with Performance Validation

Cliente Freelancer · Remoto · Remoto · freelance · mid · 750–1500 USD

Publicada el 2026-07-28

Descripción de la oferta

PROJECT SUMMARY I need an experienced developer to complement and correct a lottery prediction and betting automation system for 7 lotteries (Spain: La Primitiva, Euromillions, El Gordo; Brazil: Mega-Sena, Lotofácil, Quina, +Milionária). The system already has a substantial codebase (backend, scraping, Selenium automation, initial ML models, interface), but it has critical structural failures that make it useless for real betting. The work consists of leveraging what has already been built and implementing the necessary corrections and completions. The goal is to fix, complete, and stabilize the system, ensuring it: Continuously learns from each historical and future draw. Produces a stable and predictable ordering of combinations, with the position difference between consecutive draws ideally ≤ 15,000 positions. Is transparent and verifiable by the user (logs, checksums, performance dashboard). Allows bet export by range using fixed and immutable IDs. Has a total learning reset button with visual progress tracking. This is not a "magical" prediction system that guarantees wins – it is a continuous optimization and predictability tool, with clear metrics for the user to track learning evolution. WHAT THE SYSTEM ALREADY HAS (AND WHAT NEEDS TO BE FIXED/COMPLETED) What Already Exists (Codebase): Python backend (Django/Flask – to be confirmed during audit). Historical data scraping from the first draw of each lottery. Combination generation (full wheel) and bet export. Selenium automation for login and bet submission (with manual confirmation). Initial Machine Learning model (Gradient Boosting). Web interface and basic dashboard. Multi-lottery support. Critical Issues to Be Fixed/Completed: IDs are not fixed and immutable – the system reorders or renumbers IDs after export, breaking traceability. No continuous learning – the model trains only once at the beginning and is never updated with new draws. Broken difference pattern – position differences between consecutive draws exploded to millions, when they should be stable (≤ 15,000). Missing feedback loop – the system does not use error (difference between predicted and actual position) to adjust ordering. Missing requested models – only Gradient Boosting was used; Random Forest, LSTM (PyTorch), and Neuro-Symbolic AI are missing. Dashboard lacks learning metrics – no accuracy, mean error, rank evolution charts, etc. No visual tracking – during historical reprocessing, no progress bar, status per draw, or estimated time. No hash validation – no checksum to ensure the downloaded file (.txt) matches the generated file (.orc). No large jump alerts – the system does not notify when the difference between draws exceeds 1 million (indicating failure or manipulation). MANDATORY REQUIREMENTS (NON-NEGOTIABLE) 1. Fixed and Immutable IDs Each combination must have a unique and permanent ID in the format: LOTTERY_DRAW_XXXX_COMBO_YYYYYY The same ID must never represent a different combination. IDs cannot be renamed, reordered, or deleted at any stage. 2. Immutable Master File After each draw, a master file must be generated and permanently saved, never altered. The user can download it at any time. 3. Continuous Learning (Incremental) The system must learn from the second draw through the most recent, accumulating knowledge with each new draw. Learning must be incremental, not one-time training. 4. Stable Difference Pattern (≤ 15,000 positions) This is the most important requirement. The system must produce an ordering where the position difference of the prize between one draw and the next is small, stable, and predictable – ideally ≤ 15,000 positions (number of bets). The dashboard must display this metric clearly. The system must alert if the difference jumps to millions (indicating failure or manipulation). 5. Requested AI Models Implement Random Forest, LSTM (PyTorch), and Neuro-Symbolic AI as an ensemble, with adaptive weights (models with better recent performance gain more influence). Keep Gradient Boosting as part of the ensemble. 6. Feedback Loop (Learning from Error) After each draw, the system must calculate the error (difference between predicted and actual position) and use that error to reorder the next master file. Errors > 15,000 positions should be penalized exponentially. 7. Hash Validation (Checksum) Generate a hash (SHA-256) of the master file at creation and store it. When the user downloads the .txt file, the system must verify if the downloaded file's hash matches the stored hash. If there is a difference, the download must be blocked and an alert generated. 8. Dashboard with Learning Metrics Model accuracy per draw (line chart). Mean error per draw (difference between predicted and actual position). Top 100/1000 tickets performance (how many hit). Prize rank evolution over time. Cross-validation (10% most recent data as hidden test set). Progress bar with status per draw and estimated time during historical reprocessing. Dynamic confidence zone (automatic suggestion of the best betting range based on recent history). Outlier alert (difference > 1 million). 9. Export by Range The user can select an ID range (e.g., 112,000 to 125,000) and export as a bet file, preserving original IDs (no renumbering). 10. Lightweight Validation Log For each draw, automatically generate a record (CSV/JSON) containing: date, prize position (rank), master file checksum, and difference metric. 11. Total Learning Reset Button that deletes all previous learning and restarts processing from the very first draw, with progress bar, status per draw, and estimated time. Must preserve fixed IDs. 12. Support for 7 Lotteries Spain (3): La Primitiva, Euromillions, El Gordo. Brazil (4): Mega-Sena, Lotofácil, Quina, +Milionária. Each lottery operates independently (IDs, master files, learning, automation). 13. Complete and Documented Source Code Delivered at project completion, with all implementations and installation/deployment instructions. WHAT THE SYSTEM SHOULD PRODUCE (VISUAL EXAMPLES) Expected Pattern (Small and Stable Differences): Date Special pos Difference vs previous 2026-04-20 120,681 – 2026-04-18 119,851 830 (OK) 2026-04-16 118,860 991 (OK) 2026-04-13 116,855 2,005 (OK) Unacceptable Pattern (Explosive Differences in Millions): Date Special pos Difference vs previous 2026-07-18 101,393,579 94 MILLION (NOT OK) 2026-07-16 123,432,112 22 MILLION (NOT OK) 2026-07-13 131,382,038 7.9 MILLION (NOT OK) ESTIMATED TIMELINE 4 to 6 weeks for delivery of a stable and validated version, based on the described scope. REQUIRED SKILLS Python (Django/Flask, Selenium, Pandas, NumPy). Machine Learning / Deep Learning: Random Forest, LSTM (PyTorch/TensorFlow), Neuro-Symbolic AI, Genetic Algorithms. Databases: PostgreSQL or MongoDB. Web development: dashboards with charts (Plotly, Matplotlib). Scraping: BeautifulSoup, Scrapy. Automation: Selenium (headless), session management, CAPTCHA handling. Version control: Git. API and microservices architecture knowledge (plus). English or Spanish (for communication and documentation). REFERENCE DOCUMENTS Upon application, you will have access to the following documents (I will send after initial contact): DEFINITIVE TECHNICAL DOCUMENT lottery 2.docx – complete technical specifications of the system. Existing source code (for analysis and reuse). Examples of historical data and difference patterns. HOW TO APPLY Submit your proposal with: Brief presentation of your experience in similar projects. Suggested approach to meet the requirements (especially the difference pattern). Estimated timeline and detailed cost breakdown. Examples of previous work with web automation and/or prediction systems. IMPORTANT NOTE This is not a "magical prediction" system – it is a continuous optimization and predictability tool based on machine learning and adaptive ordering. Promises of guaranteed results will not be accepted. The focus is transparency, validation, and stability, allowing the user to track the system's learning and make informed decisions.

Skills

Fuente original: freelancer

Análisis JobHunter