ES / ENEspañolGet in touch
Work / Agent observability

Agent observability

Tracing and evaluation for the agents I run.

In progressData and observability

What it is

Phased lab for agentic-system observability, deployed on a NAS with real agents and fail-closed redaction of sensitive data. LangFuse traces every agent I run and those traces feed task scores and custom evaluators, so a harness change is judged on measured outcomes. Separately, a self-hosted Arize Phoenix traces the receipt-scan LLM calls in Tally.

Status

Done

  • OTel traces, Collector, metrics and dashboards
  • GenAI conventions, routing, retries and subagents
  • Phases 8–11: MCP and async jobs, Langfuse sessions, privacy and failure debugging
  • NAS deployment with real agents and all four runtimes through its gateway
  • SLI/SLO with dashboard and alerts to Telegram
  • coding-v1 benchmark suite: 10 tasks with results joined against telemetry

Now

Nothing in progress.

Recent changes

Summary of merged pull requests, updated daily.

  1. Telemetry-chain alerts delivered to Telegram, with a daily heartbeat
  2. coding-v1 benchmark suite: 10 tasks scored by joining telemetry into the results
  3. Session telemetry is attributed to its task run, and every runtime is identified
  4. All four runtimes send telemetry through the NAS gateway; direct ingestion retired
  5. Production SLI/SLO: four metrics, Grafana dashboard and multi-tenant guide
  6. Phase 11: failure injection and signal-driven debugging, verified on the NAS
  7. Phases 8–10: async jobs and MCP, sessions scored in Langfuse, sampling and privacy
PreviousnormaNextOAK