MLOps Expert for AI Enhancement
Publicada el 2026-07-29
Descripción de la oferta
The goal is to enhance our existing AI system so it operates as a truly production-grade platform. Today the core models are live, but our pipelines and governance layers lag behind the growing usage of generative AI. I need someone who has already shipped real-world solutions across the full stack—model training, evaluation, deployment, monitoring, and continuous improvement—and can drop straight into an AWS SageMaker environment. Key focus areas • Model lifecycle: automate versioning, lineage, and rollback using GitLab CI/CD and Prefect orchestration. • Quality & evaluation: set up repeatable LLM evaluation suites, RAG benchmarks, and human-in-the-loop (HITL) review loops so we know precisely how each release performs. • Data layer: tighten integration between Snowflake and SageMaker feature stores while enforcing governance and observability best practices. • Runtime reliability: implement comprehensive monitoring (latency, cost, drift) and alerting, surfacing metrics via built-in SageMaker tools or open-source alternatives. I am comfortable iterating in weekly milestones; each milestone should ship demonstrable value—whether that is an automated training pipeline, a new evaluation harness, or a dashboard proving lineage coverage. You will have direct access to existing repos, AWS accounts, and data warehouse connections, and can propose additional tooling if justified. Please respond with one or two concrete examples where you improved a production AI or LLM system end-to-end, the stack you used, and how you measured success. Code samples or links to public talks/posts are a plus. Looking forward to collaborating on a rock-solid, observable, and easily deployable ML/LLM platform.
Skills
Fuente original: freelancer