Skip to content

Open source · Built in Rust · OpenTelemetry traces and logs

Every failing request, explained.

Tayga turns OpenTelemetry traces and logs into error stories. For each failing or slow request it shows the root-cause span, the path across services, the critical path, what differs from the endpoint's normal, and the logs that matter. One problem is one story group, not hundreds of traces.

  • OTLP over gRPC and HTTP
  • Redpanda / Kafka API
  • ClickHouse storage
  • Self-hosted, AGPLv3 core
A Tayga error story in dark theme: the summary "payment charge failed", the root cause in the payment service's charge span, the request path from load-generator through checkout to payment, and the waterfall of the failing spans ending in the highlighted root-cause span.
An error story in the Tayga web app, captured from a running stack with the OpenTelemetry demo.

What you get

From raw spans to an explanation

Tayga does the reading for you. It assembles every trace, analyses the failing and slow ones, mines the logs, and shows the result in one web app.

  • Error stories

    The root-cause span, picked from the failing leaves and explained in one sentence, plus the request path across services and the spans that also failed.

  • Slow stories and the critical path

    A request slower than its endpoint's normal, max(p99 × 1.5, p99 + 100 ms), becomes a slow story with the critical path and its top contributors.

  • Compared with normal

    Each story is diffed against the endpoint's baseline from the last hour: new operations, missing operations and slower operations.

  • Story groups

    A fingerprint of kind, endpoint, root-cause span and masked message folds one underlying problem into one group, with its trend.

  • Log templates

    Drain mines every service’s logs into templates as they stream in. Numbers, UUIDs and hex runs are masked; HTTP status codes in access logs are kept.

  • Log alerts

    New-template and rate-spike alerts, plus opt-in silence alerts per template, delivered to webhook and Slack targets with retries.

  • Service map

    Services and their calls with rate, error ratio and p99 against the service's 24-hour baseline. Infrastructure services hide behind one switch.

  • Trace explorer

    Filter by service, endpoint, duration and errors, brush the duration scatter, and open any trace as a waterfall with a span drawer.

  • Pipeline health built in

    Component status, metric history and consumer lag, recorded by the API itself. Prometheus and Grafana stay optional.

How it works

A streaming pipeline, not a query you have to write

Tayga analyses traces as they arrive. By the time you open the app, the failing requests are already stories.

The Tayga pipelineAn OpenTelemetry Collector sends traces and logs over OTLP to tayga-ingest. Ingest publishes them to the Redpanda topic tayga.signals, keyed by trace id, and also publishes logs to tayga.logs, keyed by service. tayga-assembler reads tayga.signals and turns traces into stories; tayga-writer stores the raw spans and logs; tayga-logminer reads tayga.logs, mines log templates and raises alerts. All three write to ClickHouse. tayga-notifier delivers alerts to webhook and Slack targets, and tayga-api serves the web app and JSON API from ClickHouse.Redpanda · Kafka APItayga.alertstayga.signalskey = trace_idtayga.logskey = serviceOTel Collectoryour servicestayga-ingestOTLP gRPC + HTTPtayga-assemblertraces → storiestayga-writerraw spans + logstayga-logminerDrain · log alertstayga-apiweb app + JSON APIClickHousestories · tracestayga-notifierwebhook · SlackThe Tayga pipelineAn OpenTelemetry Collector sends traces and logs over OTLP to tayga-ingest. Ingest publishes them to the Redpanda topic tayga.signals, keyed by trace id, and also publishes logs to tayga.logs, keyed by service. tayga-assembler reads tayga.signals and turns traces into stories; tayga-writer stores the raw spans and logs; tayga-logminer reads tayga.logs, mines log templates and raises alerts. All three write to ClickHouse. tayga-notifier delivers alerts to webhook and Slack targets, and tayga-api serves the web app and JSON API from ClickHouse.Redpandatayga.signalskey = trace_idtayga.logskey = serviceOTel Collectoryour servicestayga-ingestOTLP gRPC + HTTPassemblerstorieswriterraw datalogminertemplatesClickHousestories · traces · logsnotifierwebhook/Slacktayga-apiweb app + JSON API
  1. Ingest

    Point an OpenTelemetry Collector at tayga‑ingest over OTLP (gRPC or HTTP). Spans are grouped into one record per trace id per export request on Redpanda, keyed by trace id; logs also go to a second topic, keyed by service.

  2. Assemble and analyse

    The assembler closes a trace after 10 s without new spans, or 60 s at most, builds the span tree, and turns failing and slow requests into stories with root cause, critical path and baseline diff.

  3. Explain and alert

    The logminer mines log templates per service and raises new, spike and silence alerts; the notifier delivers them. tayga‑api serves it all from ClickHouse.

Read the architecture

Tour

See a failure become a story

A narrated walk through the app: an injected failure, the story it produces, the service map, log templates and alerts, and the pipeline.

Benchmarks

Measured, dated, sourced

Numbers from the performance report, each with its method and source. Nothing is extrapolated.

  • 2.8µs

    per log line through Drain

    drain_add over a 50,000-line corpus from the stack, one thread

  • 6.98×

    faster mining with the fingerprint cache

    cached/add vs add_fingerprinted, same corpus

  • 99.70%

    cache hits on the live stack

    69,359 of 69,569 lines, first 30 minutes after deploy

  • 1.07%

    of one core for the logminer

    live demo load, about 40 log lines per second

Criterion micro-benchmarks in a release build on an Apple M3 Max (14 cores); live figures from the OpenTelemetry demo stack on 6 and 7 October 2026. Method and full results

Comparison

Where Tayga fits, honestly

Tayga is a focused tool: it explains failing and slow requests. Here is what it is built for, and when another tool is the better choice.

Choose Tayga for

  • Per-request stories, built while traces stream in. Every failing or slow request gets an error or slow story as its trace closes, not when someone writes a query.
  • Deterministic rules, no LLM. Root cause, critical path and baselines follow rules you can read. The same input gives the same answer.
  • Plain OpenTelemetry in. Anything that speaks OTLP, directly or through a Collector. No proprietary agent.
  • Open source, log alerts in the core. New, spike and opt-in silence alerts on log templates are in the AGPL core, self-hosted: your telemetry stays in your ClickHouse and Redpanda.

Look elsewhere if

  • You need metrics. Tayga takes traces and logs only. Keep a metrics backend next to it.
  • You need full-text log search. Tayga mines logs into templates and links them to stories; it is not a log search engine.
  • You want a managed service. The open-source edition is self-hosted.

The full comparison covers, with a source and access date for every cell:

  • Jaeger
  • Grafana Tempo + Loki
  • SigNoz
  • Coroot
  • OpenObserve
  • ClickStack / HyperDX
  • Datadog
  • Dynatrace
  • New Relic
Read the comparison

Tayga Enterprise · available on request

Running Tayga across teams and clusters?

The enterprise edition adds identity, isolation, scale and support on top of the open-source core, under a commercial license. It is priced per cluster or node, not per GB ingested.

What the enterprise edition includes

  • SSO / SAML, SCIM, fine-grained RBAC and audit logs
  • Multi-tenancy and multi-cluster federation
  • High availability, scale-out deployment and upgrade tooling
  • PII redaction policies and data-residency controls
  • Long-term baselines: compare a release with previous deploys
  • ServiceNow, Jira and PagerDuty integrations
  • LLM incident summaries with your own model
  • Support with an SLA, onboarding and training

Give every failing request a story

Run Tayga next to the OpenTelemetry demo on your laptop, or point your own Collector at it.

Named after Tajga, a hunting dog who can trace anything, anywhere.Why the name