Work / Agent observability
Agent observability
Tracing and evaluation for the agents I run.
What it is
Phased lab for agentic-system observability, deployed on a NAS with real agents and fail-closed redaction of sensitive data. LangFuse traces every agent I run and those traces feed task scores and custom evaluators, so a harness change is judged on measured outcomes. Separately, a self-hosted Arize Phoenix traces the receipt-scan LLM calls in Tally.
Status
Done
- OTel traces, Collector, metrics and dashboards
- GenAI conventions, routing, retries and subagents
- Phases 8–11: MCP and async jobs, Langfuse sessions, privacy and failure debugging
- NAS deployment with real agents and all four runtimes through its gateway
- SLI/SLO with dashboard and alerts to Telegram
- coding-v1 benchmark suite: 10 tasks with results joined against telemetry
Now
Nothing in progress.
Later
No firm plans.
Recent changes
Summary of merged pull requests, updated daily.
- Telemetry-chain alerts delivered to Telegram, with a daily heartbeat
- coding-v1 benchmark suite: 10 tasks scored by joining telemetry into the results
- Session telemetry is attributed to its task run, and every runtime is identified
- All four runtimes send telemetry through the NAS gateway; direct ingestion retired
- Production SLI/SLO: four metrics, Grafana dashboard and multi-tenant guide
- Phase 11: failure injection and signal-driven debugging, verified on the NAS
- Phases 8–10: async jobs and MCP, sessions scored in Langfuse, sampling and privacy