SRE Guidance: Monitoring and Performance
Publicada el 2026-07-23
Descripción de la oferta
I am handling the day-to-day reliability of a growing production stack and need a seasoned Site Reliability Engineer to coach me through two key domains: monitoring/alerting and performance optimization. I already manage the basic upkeep, but I want to elevate my skills so that incidents are caught sooner and services run leaner. Here is what I’m hoping for: • Regular screen-sharing or video sessions where we review my existing dashboards and alert rules, refine the signal-to-noise ratio, and discuss industry best practices. • Deep-dive walkthroughs on performance tuning—profiling services, interpreting latency metrics, and translating findings into configuration or code changes. • Actionable take-home steps after each session so I can apply what we discuss, then bring results back for feedback. I work primarily in a Linux/containerized environment with common open-source tooling, but I’m open to adopting whatever stack you recommend—Prometheus, Datadog, Grafana, or other fit-for-purpose solutions. To make sure we are a match, please tell me about: • Similar mentoring or advisory roles you’ve done. • Your approach to setting up meaningful alerts without alert fatigue. • A brief example of how you diagnosed and fixed a tricky performance bottleneck. We can start with a short engagement to align on goals; if the collaboration clicks, I’m happy to extend for ongoing guidance.
Skills
Fuente original: freelancer