Architecture
Tayga is a streaming pipeline. Telemetry enters through one OTLP endpoint, travels on Redpanda topics, and is analysed by independent services while it streams. Each service is a separate Rust binary from one workspace; they exchange data only through the bus and ClickHouse (the API also scrapes their metrics), so each can be restarted, upgraded or scaled on its own.
The flow
Section titled “The flow”- Ingest. An OpenTelemetry Collector (or any OTLP sender) exports traces and logs to
tayga-ingest, over gRPC or HTTP. For each export request, ingest publishes one record per trace id to the topictayga.signals, keyed by trace id, so every span of a trace lands on the same partition (a record is split further if it would pass 900,000 bytes; items without a valid trace id are keyed by service). It also publishes each log batch totayga.logs, keyed by service, one record per request and service. - Store.
tayga-writerreadstayga.signalsand writes the raw spans and logs to ClickHouse. - Assemble and analyse.
tayga-assemblerreads the same topic, buffers spans per trace, and closes a trace when it goes quiet. For each closed trace it builds the span tree, writes a trace summary, counts service-to-service calls, and, when the request failed or was unusually slow, writes an error story to ClickHouse and publishes it totayga.stories. - Mine the logs.
tayga-logminerreadstayga.logs, mines templates per service with Drain, and every 60 seconds checks for new, spiking and silent templates. Alerts go to ClickHouse and to the topictayga.alerts. - Deliver.
tayga-notifierreadstayga.alertsand delivers each alert to webhook and Slack targets. - Serve.
tayga-apiserves the web app and the JSON API from ClickHouse, records the services’ metrics for the Pipeline page, and reads consumer lag from Kafka.
Services
Section titled “Services”| Service | Reads | Writes | Role |
|---|---|---|---|
tayga-ingest |
OTLP gRPC and HTTP | tayga.signals, tayga.logs |
Accepts traces and logs, splits them by trace id and by service |
tayga-writer |
tayga.signals |
ClickHouse spans, logs |
Stores raw telemetry; also runs the schema migrations (tayga-writer migrate) |
tayga-assembler |
tayga.signals |
trace_summaries, service_edges, error_stories; tayga.stories |
Closes traces, analyses them and writes stories |
tayga-logminer |
tayga.logs |
log_templates, log_template_hits, log_alerts; tayga.alerts |
Mines log templates and raises log alerts; can run as several replicas |
tayga-notifier |
tayga.alerts |
notifier_deliveries; webhook and Slack targets |
Delivers alerts once per target, with retries |
tayga-api |
ClickHouse, Kafka consumer lag, the services’ /metrics |
metric_samples, log_template_silence |
Web app, JSON API, metric recorder |
The web app is a React single-page app built into the tayga-api binary, so one port serves both the app and the API.
Topics
Section titled “Topics”| Topic | Key | Partitions | Producer | Consumers |
|---|---|---|---|---|
tayga.signals |
trace id | 12 | ingest | writer, assembler |
tayga.logs |
service | 12 | ingest | logminer |
tayga.stories |
story fingerprint | 3 | assembler | none in Tayga; for other consumers |
tayga.alerts |
template id | 3 | logminer | notifier |
Partition counts apply when a topic is created; see Replicas and partitioning. Every log is written to Kafka twice, once per topic; spans are not duplicated. Topics that Tayga creates get a retention of 24 hours by default.
Closing a trace
Section titled “Closing a trace”The assembler keeps the open traces of each partition in memory and closes a trace on the first of these conditions, checked every second:
| Condition | Default | Effect |
|---|---|---|
| No new span for | 10 s (processing time) | closed normally |
| Open for longer than | 60 s | closed normally |
| Span count above | 10,000 | closed, flagged truncated |
| All buffered traces above | 512 MB | the oldest are closed, flagged truncated |
Inactivity uses processing time because clock skew between services makes event-time watermarks unreliable; the analysis itself uses the spans’ own timestamps. Spans that arrive after their trace closed are counted and stored raw, but not re-analysed.
Analysis
Section titled “Analysis”The analysis code (tayga-analysis) is pure functions over a trace and the endpoint’s baseline, with no I/O:
- Span tree. Built from parent span ids; logs attach to their span, or to the root when they carry only a trace id. The endpoint is the root service and the root span’s name, with a normalised path added when the name has none.
- Root cause. The error span among those with no failing descendant that ended first, with a one-sentence explanation. See Root cause and critical path.
- Critical path. A backward walk from the root span’s end, with the top three contributors by self-time.
- Baseline diff. Per endpoint, refreshed every 60 seconds from the last 60 minutes of non-error traces. See Baselines.
- Story and fingerprint.
errorwhen any span failed,slowwhen the duration exceeded the slow-trace rule on a trusted baseline. The fingerprint groups stories into story groups.
Storage
Section titled “Storage”| Table | Written by | Keeps |
|---|---|---|
spans, logs |
writer | 3 days |
trace_summaries |
assembler | 2 days |
error_stories, service_edges |
assembler | 7 days |
log_template_hits |
logminer | 3 days |
log_templates |
logminer | 30 days after the template was last seen |
log_alerts |
logminer | 7 days |
notifier_deliveries |
notifier | 30 days |
metric_samples |
api | 7 days |
Delivery and replays
Section titled “Delivery and replays”Offsets are committed only after the data they cover is stored, so a crash means re-reading, never losing. Re-reads are safe by design:
- stories use the trace id as their id, and stories and trace summaries are stored in
ReplacingMergeTreetables, so a replayed trace collapses into one row; - log hits are deduplicated by log id, and template ids are hashes of service and template;
- alert ids are deterministic, and the notifier records each delivery per alert and target, so a re-published alert is not sent twice.
Workspace
Section titled “Workspace”| Crate | Kind |
|---|---|
tayga-ingest, tayga-writer, tayga-assembler, tayga-logminer, tayga-notifier, tayga-api |
services |
tayga-analysis |
the story analysis, no I/O |
tayga-drain |
Drain mining and alert rules, no I/O |
tayga-model, tayga-kafka, tayga-store, tayga-common |
shared libraries |
tayga-devtools |
CLI: flags, capture, emit-log, hash-password, remine and more |
tayga-e2e |
end-to-end tests against a live stack |
