Kurdish LLM Fine-Tuning Advisor
Publicada el 2026-07-19
Descripción de la oferta
I’m building a production-grade large-language model for Kurdish, covering both Sorani and Badini, and I need an experienced fine-tuner to guide the technical core of the project. The first milestone is our data pipeline. The dataset structure is critical: I want clearly separated CPT and SFT splits, stored as clean JSONL, with solid deduplication and quality filters baked in rather than patched on later. You’ll help me design that pipeline end-to-end instead of just running someone else’s script. Next comes training. We are using QLoRA on top of Hugging Face Transformers with PEFT and, ideally, Unsloth for speed-ups. I need hands-on advice across the entire stack—model configuration, training optimisation and hyper-parameter tuning—so my in-house engineer can execute confidently. After the initial setup I’ll book you for 10-20 hours each month for reviews, troubleshooting and iteration planning. Preference goes to someone who has already fine-tuned and published non-English models (please link your Hugging Face profile) and who understands Arabic-script or other low-resource languages. Deliverables • Detailed data-pipeline specification with scripts or notebooks • QLoRA training config (model, optimiser, scheduler, eval) plus rationale • Ongoing written or call-based guidance, logged as actionable notes for our engineer If you have the background, especially in low-resource or Arabic-script NLP, let’s talk.
Skills
Fuente original: freelancer