
Welcome to Observability 101, a hands-on course about turning production behavior into numbers you can query, graph, alert on, and trust.
If the answer to any of these questions is yes, then this course is for you.
The course aims to teach you the fundamentals of observability and fill in the gaps people often have when starting out.
The course starts with the big picture: telemetry, monitoring, observability, metrics, logs, traces, and profiles. After that, it narrows to metrics. Metrics are not the whole observability story, but they are usually the highest-value first signal: cheap to collect, compact to store, fast to query, and good enough to power most dashboards, alerts, and SLOs.
Before starting, you should be comfortable with:
curl, and reading JSON/YAML.No Prometheus, VictoriaMetrics, OpenTelemetry, Kubernetes, or SRE background is assumed.
By the end of this course, you'll be able to:
The course is organized into nine modules that build on each other, followed by a capstone:
| Module | Focus | Lessons | Status |
|---|---|---|---|
| 1. Welcome | Core vocabulary, telemetry signals, why metrics come first, and how this course teaches | 2 | Available |
| 2. Metrics data model | Names, labels, samples, series, cardinality, and metric types (counters, gauges, histograms, summaries) | 3 | Available |
| 3. Collection and storage | How metrics move from process to storage: exposition format, scraping targets, pushing, vmagent, and storage with VictoriaMetrics | 5 | Coming soon |
| 4. Querying | PromQL/MetricsQL selectors, rates, ratios, and histogram percentiles that answer operational questions | 5 | Coming soon |
| 5. Visualization | Dashboards in vmui and Grafana, and a repeatable investigation loop | 3 | Coming soon |
| 6. Application instrumentation | Deciding what your service should measure (service signals), then designing and exposing safe metrics and labels from your own code | 4 | Coming soon |
| 7. Infrastructure monitoring | The metrics you get without writing code: host and container exporters, plus external probes of user journeys | 2 | Coming soon |
| 8. Alerting | Alerting with vmalert and Alertmanager, plus SLIs and SLOs | 3 | Coming soon |
| 9. Pipeline reliability | Keeping a metrics pipeline reliable: cardinality limits, relabeling, dropping, retention, downsampling, deduplication, and HA | 4 | Coming soon |
A single hands-on scenario that ties everything together: metrics-driven incident response, where you use what you've learned to investigate and resolve a production issue end to end. Coming soon, once the modules it draws on are in place.
Ready to start?
Writes about
Frequently covers