Skip to content
Back to skills

Observability

ASecurity

Instrument Python services and agents with structured logs, traces, and metrics.

  • 10 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 23, 2026
ai-agentspythongobashfastapigitapibackenddocumentation

Works with

  • cli
  • api

Security analysis

A100/100

Scanned October 6, 2026

npx -y skills add fmind/dot --skill observability --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Observability?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Observability
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fmind-observability/badge)](https://www.skillsdirectory.com/skills/fmind-observability)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: observability
description: "Instrument Python services and agents with structured logs, traces, and metrics."
license: MIT
metadata:
  kind: task
  author: Médéric HURIER (Fmind)
  source: github.com/fmind/dot/tree/main/skills/observability
  created: "2026-09-03"
  updated: "2026-10-05"
---

# Observability

Use one Python telemetry stack for services and agents: `structlog` JSON on stdout, OpenTelemetry traces and metrics over OTLP, and trace context on every related log event. Standard OTLP configuration keeps the application independent of the backend.

## Workflow

1. **Add only the Python packages in use**: `structlog`, OpenTelemetry API and SDK, the OTLP exporter, and explicit instrumentation packages for the service's HTTP framework and clients; FastAPI 0.142+ instruments itself ([fastapi](../python-web/references/fastapi.md)). Lock them with `uv`; avoid a vendor SDK in application code.
1. **Emit structured logs**: render one JSON object per line to stdout in production. Use `severity`, `message`, and `time` for Cloud Logging while retaining stable event names and machine-readable fields.
1. **Configure traces and metrics**: use `OTEL_SERVICE_NAME`, `OTEL_RESOURCE_ATTRIBUTES`, `OTEL_EXPORTER_OTLP_ENDPOINT`, and an explicit protocol. Keep local development quiet by disabling export when it is not configured; an SDK's default localhost endpoint does not disable telemetry.
1. **Correlate signals**: a `structlog` processor reads `trace.get_current_span().get_span_context()` and adds `trace_id`, `span_id`, `logging.googleapis.com/trace`, and `logging.googleapis.com/spanId` only when the context is valid.
1. **Describe agent work**: use current GenAI semantic conventions; model calls carry `gen_ai.operation.name`, provider and request model, and input/output token usage; agent and tool spans carry their stable agent or tool names. Do not record prompt or completion bodies by default.
1. **Export through a collector sidecar**: on Google Cloud, use the Google-built OpenTelemetry Collector as a Cloud Run sidecar. Match the exporter and receiver: `grpc` with `http://localhost:4317`, or `http/protobuf` with `http://localhost:4318`. For environment-configured instrumentation set `OTEL_EXPORTER_OTLP_PROTOCOL`; FastAPI's automatic export requires `http/protobuf`. For manually constructed Python exporters choose the matching gRPC or HTTP class. Let the collector authenticate with ADC and forward telemetry to Google Cloud.
1. **Evaluate separately**: operational telemetry detects failures and drift but does not prove response quality. Link a trace ID to Langfuse or MLflow scores when used, and use [agent-evaluation](../agent-evaluation/SKILL.md) for repeated comparisons through the project's existing runner.
1. **Verify all three signals**: send one request, locate its trace, read the correlated log events, confirm the expected metric, then exercise shutdown to prove buffered telemetry flushes within the platform grace period.
   ```bash
   gcloud logging read 'trace="projects/<project>/traces/<trace_id>"' --project=<project> --freshness=1h --limit=20 --format='json(timestamp,severity,textPayload,jsonPayload.message,trace,spanId)'
   ```

## Gotchas

- **Bound investigation output**: select incident time and correlation IDs before loading logs; include custom event fields when needed for diagnosis. A limit is a partial view, not proof of absence. Keep complete authorized captures as local artifacts and expand the query when the cause falls outside the initial window.
- **Cloud Logging keys are exact**: plain `level` and `msg` are not promoted to severity and message fields.
- **Sampling follows the parent**: use `parentbased_traceidratio` with a measured production ratio; keep 100% sampling for bounded development only.
- **Cardinality is a budget**: do not put user IDs, raw URL IDs, prompts, errors, or tool arguments in metric labels.
- **Telemetry is best effort**: bound exporter queues and timeouts, handle `SIGTERM`, and never block a user request on an unavailable collector.
- **Privacy starts before export**: redact secrets and personal data in the application, because backend filters cannot retract already-exported spans or logs.

## Official Skills

Upstream: `langfuse/skills`, `mlflow/skills`, `pydantic/skills`, and `grafana/skills`. Follow the shared [vendor-skill policy](../agent-project/references/vendor-skills.md) and install only the backend selection the project uses. The Vercel plugin in `openai/plugins` ships an unrelated same-name `observability` skill; never install it beside this one. Exclude `mlflow/skills`' same-name `agent-evaluation` from any selection ([name policy](../agent-project/references/vendor-skills.md#name-collisions)).

## Documentation

- [OpenTelemetry Python](https://opentelemetry.io/docs/languages/python/) · [OTLP configuration](https://opentelemetry.io/docs/languages/sdk-configuration/otlp-exporter/) · [GenAI semantic conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/) · [Cloud Logging structured logs](https://docs.cloud.google.com/logging/docs/structured-logging) · [Google-built OTel Collector](https://docs.cloud.google.com/stackdriver/docs/instrumentation/google-built-otel)
- Releases: [OpenTelemetry Python](https://github.com/open-telemetry/opentelemetry-python/releases)
- Companion skills: [python-stack](../python-stack/references/foundation/GUIDE.md), [quality-assurance](../quality-assurance/SKILL.md), [cloud-run](../cloud-run/SKILL.md), [google-adk](../agent-frameworks/references/google-adk.md), [gcloud](../gcloud/SKILL.md), [benchmark](../benchmark/SKILL.md).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…