Proficient AI Data Generation and Evaluation

Cliente Freelancer · Remoto · Remoto · freelance · mid

Publicada el 2026-07-19

Descripción de la oferta

I have worked on multiple AI data generation and model evaluation projects involving both Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF). My work included designing high-quality prompts, analyzing model responses, identifying reasoning and factual errors, and creating golden responses that served as the expected reference outputs. I worked extensively with repository-based tasks using Cursor, where I understood large codebases, executed run scripts, analyzed parse scripts, and validated both Pass-to-Pass (P2P) and Fail-to-Pass (F2P) test cases to ensure the generated solutions were functionally correct. I also reviewed code quality, verified outputs against expected behavior, and ensured compatibility with existing repositories. For tool-use and agentic AI projects, I created prompts to evaluate model capabilities, intentionally designed scenarios to expose model weaknesses, and assessed how effectively models selected and invoked the available tools. I wrote detailed evaluation rubrics defining the expected reasoning process, required tool schemas, correct tool calls, expected outputs, and grading criteria. I compared multiple model responses for correctness, instruction following, reasoning quality, completeness, and safety, while identifying issues such as hallucinations, incorrect tool usage, missing steps, and formatting problems. Throughout these projects, I followed strict quality guidelines, produced consistent annotations, documented evaluation decisions, and collaborated through review workflows to improve dataset quality and model performance. This experience has given me strong expertise in AI data annotation, prompt engineering, rubric writing, code evaluation, tool-use assessment, and LLM response evaluation.

Skills

Fuente original: freelancer