AI Evaluator & Data Annotator Needed

Cliente Freelancer · Remoto · Remoto · freelance · mid · 250–750 CAD

Publicada el 2026-07-21

Descripción de la oferta

## AI Evaluation & Data Annotation Specialist — LLM, Python, Linux and Multimodal Tasks I am looking for a talented and detail-oriented **AI Evaluation and Data Annotation Specialist** for ongoing work across multiple AI training, benchmarking, and evaluation projects. The projects may include technical LLM evaluation, terminal-based tasks, coding-response verification, speech transcription, audio/video annotation, image and text evaluation, ranking tasks, prompt assessment, and quality assurance. ### Responsibilities * Evaluate LLM and AI-generated responses for correctness, relevance, completeness, and quality * Create, test, and validate AI benchmark tasks * Work with Linux terminals, command-line tools, Python, Bash, Git, and Docker * Test and verify AI-generated code or technical solutions * Identify hallucinations, factual errors, missing details, and inconsistencies * Transcribe speech accurately * Annotate text, audio, images, and videos * Add accurate timestamps and segment speech or visual events * Rank AI responses, search results, or other generated outputs * Review annotations and perform quality assurance * Follow detailed project-specific guidelines precisely * Learn new evaluation and annotation platforms quickly ### Required skills * Experience with AI evaluation, LLM benchmarking, data annotation, or AI training-data projects * Strong Linux and command-line knowledge * Working knowledge of Python and Bash * Ability to test and debug technical solutions * Experience with text, audio, image, or video annotation * Strong written English and listening comprehension * Excellent attention to detail * Ability to understand and consistently apply complex instructions * Reliable communication and availability for ongoing work ### Preferred experience * Evaluating generative AI or LLM outputs * Coding benchmarks or terminal-based evaluation tasks * Prompt engineering, rubric creation, or test-case development * SuperAnnotate, Label Studio, Scale AI, Appen, Remotasks, or similar platforms * Multimodal AI projects involving text, audio, images, and video * Git, Docker, APIs, JSON, and software testing * Annotation reviewing or quality-assurance work ### Application questions Please include the following information in your proposal: 1. Describe your experience with AI evaluation, LLM benchmarking, or data annotation. 2. What experience do you have with Linux, Python, Bash, Git, and Docker? 3. Have you evaluated AI-generated code or technical responses? 4. Have you completed text, speech, audio, image, or video annotation projects? 5. Which AI evaluation or annotation platforms have you used? 6. How do you verify that an AI-generated response is correct? 7. Are you comfortable switching between different project types and guidelines? 8. Please provide a relevant work sample with confidential information removed. ### Engagement details * Remote freelance contract * Part-time initially * Potential for ongoing work * Variable workload depending on project availability * You can get paid by 40% of each task budget.

Skills

Fuente original: freelancer

Análisis JobHunter