Verified claims
The docs make claims about what Tayga does, how it behaves and how fast it is. This page lists them with their evidence and the date each was checked. The rule behind it: a claim that cannot be checked is left out, not guessed.
Evidence is one of:
- code: read in the source (a setting’s default, a constant, a query), with the file named;
- test: a unit, integration or end-to-end test that pins the behaviour;
- live: observed on a running stack (the OpenTelemetry demo 3.1.0 with Tayga, on an Apple M3 Max with Docker Desktop), with the time;
- source: a vendor’s own page, for the comparison, accessed 2026-10-07;
- commercial offering (owner): what the owner states about the enterprise edition. These are not features of the open-source code and are not verified in it.
A single live run is a single run: timings below are observations, not guarantees. If you find a claim that is wrong, open an issue on GitHub or write to hello@softberries.dev.
Claims on this site
Section titled “Claims on this site”Checked on 2026-10-07 against the code on branch feat/docs-launch and the live API, unless the row says otherwise.
Landing page
Section titled “Landing page”| Claim | Evidence | Result |
|---|---|---|
| OTLP over gRPC and HTTP | code: crates/tayga-ingest/src/main.rs (trace and log gRPC services), http.rs (/v1/traces, /v1/logs) |
verified |
| Redpanda / Kafka API, ClickHouse storage | code: tayga-kafka (rdkafka), tayga-store; both Compose bundles |
verified |
| Self-hosted, AGPLv3 core | LICENSE (GNU AGPL v3), license = "AGPL-3.0-only" in Cargo.toml |
verified |
| A story shows the root-cause span, the path, the critical path, what differs from normal, and the logs | code: Story in crates/tayga-analysis/src/story.rs; live GET /api/v1/stories/2b061ed5… |
verified |
Slow rule max(p99 × 1.5, p99 + 100 ms) |
code: Thresholds::default, slow_limit_ns in crates/tayga-analysis/src/baseline.rs |
verified |
| Baseline from the last hour; new, missing and slower operations | code: baseline_window_minutes: 60; diff in baseline.rs |
verified |
| Fingerprint of kind, endpoint, root-cause span and masked message | code: build_story and fingerprint in tayga-analysis |
verified |
| Numbers, UUIDs and hex runs masked; HTTP status codes kept | code: mask (fingerprint.rs), tokens (crates/tayga-drain/src/preprocess.rs); unit tests keeps_status_after_http_version, other_numbers_stay_masked |
verified |
| New, spike and opt-in silence alerts, to webhook and Slack with retries | code: crates/tayga-drain/src/detect.rs, crates/tayga-notifier; live make e2e-notifier 2026-10-06 |
verified |
| Service map: rate, error ratio and p99 against a 24-hour baseline; infrastructure behind one switch | code: health, HEALTH_BASELINE_SECS in crates/tayga-api/src/params.rs; [map] infra_services; live GET /api/v1/service-map |
verified |
| Trace explorer: filters, duration scatter with brush, waterfall and span drawer | the user guide’s screenshots (2026-10-07); traces/search parameters in params.rs |
verified in the app and code |
| Pipeline health recorded by the API itself | code: crates/tayga-api/src/recorder.rs; live Pipeline page |
verified |
| One record per trace id per export request on Redpanda, keyed by trace id; logs keyed by service (How it works) | code: crates/tayga-ingest/src/records.rs: one record per trace id per export request, split above 900,000 bytes; items without a trace id are keyed by service |
verified; fixed: an earlier wording, “each trace becomes one record”, was imprecise, and the README, the landing page and Architecture now state this rule |
| A trace closes after 10 s without new spans, or 60 s at most | code: AssemblerSettings::default (gap_ms 10000, max_age_ms 60000) |
verified |
| 2.8 µs per log line through Drain (2.818) | docs/perf/sp4-performance.md, stages/drain_add |
verified (measured 2026-10-06/07) |
| 6.98× faster mining with the fingerprint cache | same report, cached/add 144.29 ms against cached/add_fingerprinted 20.678 ms |
verified |
| 99.70 % cache hits on the live stack (69,359 of 69,569) | same report, live 2026-10-07 04:30–05:01 UTC | verified |
| 1.07 % of one core for the logminer at about 40 lines/s | same report, docker stats 2026-10-06 18:30–19:00 UTC |
verified |
| Apple M3 Max, 14 cores; live figures from 6 and 7 October 2026 | same report, hardware table | verified |
| Tayga takes traces and logs only; no proprietary agent | code: ingest registers only the trace and log services | verified |
| Enterprise: SSO/SAML, SCIM, RBAC, audit logs; multi-tenancy and federation; HA and upgrade tooling; PII redaction and data residency; long-term baselines; ServiceNow, Jira, PagerDuty; LLM summaries; SLA support | owner | commercial offering (owner), not in the open-source code |
| Enterprise “available on request”, priced per cluster or node, not per GB | owner | commercial offering (owner) |
| The name note | owner | owner’s words |
Concepts
Section titled “Concepts”| Claim | Evidence | Result |
|---|---|---|
Error span: status error, an exception event, or an attached log with severity ≥ 17 |
code: error_flags in rootcause.rs, SEVERITY_ERROR = 17 |
verified |
Root cause = the error leaf that ended first; other leaves are also_failed |
code: find_root_cause; unit tests in rootcause.rs |
verified |
Message order: status message, exception.message, first ERROR log, error |
code: error_message |
verified |
| “could not reach” sentence for client spans without children; peer and operation attribute order | code: explain, peer_of, operation_of; live group summaries such as checkout could not reach oteldemo.PaymentService (oteldemo.PaymentService/Charge): … |
verified |
| Critical path: backward walk, 5 ms clock-skew tolerance, top 3 contributors | code: critical_path.rs (CLOCK_SKEW_TOLERANCE_NS), TOP_CONTRIBUTORS = 3 |
verified |
| A live story: 140 spans, 186 segments, top contributor 4.63 ms | live GET /api/v1/stories/2b061ed5d124c26643e501e934cb4267, 2026-10-07 |
verified |
| Story logs: at most 50, most severe first | code: MAX_STORY_LOGS = 50 and the sort in build_story |
verified |
Flags incomplete and truncated |
code: FLAG_INCOMPLETE (tree.rs), FLAG_TRUNCATED (crates/tayga-assembler/src/window.rs) |
verified |
| Late spans counted, stored raw, not re-analysed; 100,000 closed ids remembered per partition | code: window.rs (recent, late_items), recent_per_partition default |
verified |
tayga.stories keyed by fingerprint, 3 partitions |
code: pipeline.rs (story_messages), stories_partitions: 3 |
verified |
Baseline defaults (60 min window, 60 s refresh, 50 traces, 1 %, 95 %, max(p95 × 2, p95 + 50 ms)) |
code: AssemblerSettings::default, Thresholds::default, diff |
verified |
| Outlier exclusion, 10 × p50 bootstrap, carry for at most 2 windows, in memory only | code: crates/tayga-assembler/src/baselines.rs, crates/tayga-store/src/store.rs; tests listed under Plan 7a below |
verified in code and tests; not run live |
| Fingerprint parts and masking (UUID, hex ≥ 8, numbers) | code: fingerprint.rs |
verified |
Two groups for one failure, one per calling endpoint (user_checkout_single, user_checkout_multi) |
live GET /api/v1/story-groups?since=1h, 2026-10-07 |
verified |
Group summary and sample are the newest story; top 100 by count; service = root-cause service |
code: GROUP_COLUMNS, FILTERED in crates/tayga-api/src/repo.rs |
verified |
Edges: one observation per span whose parent is in another service; per minute, SummingMergeTree, 7 days |
code: service_edges in summary.rs; migration 0002 |
verified |
| Health: error at ≥ 5 % failed server/consumer spans, slow above 2 × the 24 h p99 | code: HEALTH_ERROR_RATIO, HEALTH_SLOW_FACTOR, health in params.rs |
verified |
| Map CPU per refresh 291.5 → 40.0 ms | docs/perf/sp4-performance.md, live 2026-10-07 |
verified |
| Drain defaults: similarity 0.5, depth 4, 100 children, 5,000 templates per service, 64 tokens | code: DrainConfig::default, MAX_TOKENS |
verified |
| 67 templates with wildcard scoring against 5,942 classic, on the prototype’s sample | measured on a 20,000-log demo sample while Drain was designed (plan 4), recorded in the module doc of crates/tayga-drain/src/drain.rs |
cited, not re-run |
| 423 templates held in memory | logminer /metrics at 2026-10-07 04:19 UTC (sub-project 4 control window) |
verified live |
Status-code rule (100–599 after an HTTP/x token) and exact matching |
code: is_status, is_http_version, tokens; tests under Plan 7a below |
verified |
| Fingerprint cache: 10,000 entries per service; resets; identical results | code: CACHE_MAX_PER_SERVICE, add_fingerprinted; tests/differential.rs |
verified |
| Template tables and TTLs (30 days, 3 days, 8 days) | migrations 0004, 0007 |
verified |
| Alert rules and defaults (5 min, 60 min, × 5, ≥ 10, warmup 15 min, active 10 min, detection 60 s) | code: DetectConfig::default, LogminerSettings::default |
verified |
| Coverage rule (half the minutes), young templates (10 to 65 min) | code: detect.rs; unit and store tests (Plan 7a rows below) |
verified in code and tests; no live skip observed |
| Seasonal mode not run live | Plan 7a rows below | stated as such |
New and silence alerts carry window_count 0 |
code: new_alert and the silence alert in detect.rs |
verified |
| Alert ids: hashes of kind and template, plus start minute or last hit | code: id_hex calls in detect.rs |
verified |
| A template that appeared while the logminer was stopped is announced after the restart | live 2026-10-04 (Plan 4 rows below) | verified (one run) |
| Data clock: per partition, held while behind (> 60 s), clamped to the wall clock, 10 min start offset | code: crates/tayga-logminer/src/main.rs, initial_watermark; tests (Plan 7a rows below) |
verified in code and tests |
| Overview data lag 0.98 s | live GET /api/v1/overview, 2026-10-07 |
verified (one reading) |
| Topic keys and partitions (12, 12, 3, 3) | code: KafkaSettings defaults, ALERTS_PARTITIONS, stories_partitions |
verified |
tayga.logs about 53 MB/h; tayga.signals 25.4 GB under 7-day retention |
live 2026-10-06 (Plan 8 rows below) | verified |
| Ownership 60 min; revoke pass 30 s; flush 15 s; shutdown pass 20 s; grace 40 s | code: LogminerSettings::default, REVOKE_PASS_TIMEOUT, SHUTDOWN_FLUSH_TIMEOUT, SHUTDOWN_PASS_TIMEOUT; Compose stop_grace_period |
verified |
| Republish: 24 h, 1,000 per pass | code: REPUBLISH_WINDOW_NS; Store::unpublished_alerts |
verified |
Alerting
Section titled “Alerting”| Claim | Evidence | Result |
|---|---|---|
Notifier defaults (public_url, kinds, 8 attempts, 10 s timeout ≤ 15, max_age_secs 3600, breaker 300 s) |
code: NotifierSettings::default, validate, TIMEOUT_MAX_SECS |
verified |
targets and kinds only in the file |
code: both are lists; load_settings has no list parsing |
verified |
| URLs never logged, labelled or stored | code: WebhookUrl (Debug/Display print <redacted>), error redaction |
verified in code |
| Webhook body fields, summary texts, at most 3 trace links | code: webhook_payload, summary, MAX_TRACE_LINKS in payload.rs; live body from make e2e-notifier 2026-10-06 |
verified |
| The spike example payloads | built by the rules of payload.rs from live alert b8b44f4f67291acc as stored at 13:34 UTC on 2026-10-07 |
derived, not a captured delivery |
| Slack blocks, escaping, 150 / 2,800 / 3,000 character limits, backtick replacement | code: slack_payload, escape, constants in payload.rs |
verified in code |
Retries, Retry-After, backoff cap 300 s, 40 min poll interval |
code: deliver.rs, MAX_POLL_INTERVAL_MS |
verified in code and tests; not with a live failing target |
| Breaker behaviour; “17 to 28 alerts” and “over 100 alerts an hour” | code and the plan 7b/8 analysis in the README history; arithmetic 3600 / (127 to 207 s) | verified in code; the alert rate observed live |
| Once per alert and target, also after a restart | live make e2e-notifier 2026-10-06 (Plan 7b rows below) |
verified (one run) |
| The three resend cases | code: deliver.rs docs; migration 0011 TTL |
verified in code; accepted, not reproduced |
Operations and reference
Section titled “Operations and reference”| Claim | Evidence | Result |
|---|---|---|
| Every setting, environment variable and default in the configuration reference | code: the settings structs and serde defaults of each crate; a script checked every field against the page, 2026-10-07 |
verified |
TAYGA_CONFIG file, then TAYGA__* overrides; lists cannot come from the environment |
code: load_settings in crates/tayga-common/src/lib.rs; live error for TAYGA__MAP__INFRA_SERVICES (Plan 6 rows) |
verified |
[kafka] validated by ingest, writer, assembler, logminer and notifier |
code: settings.kafka.validate() in those five main.rs |
verified |
| Authentication behaviour | Plan 6 rows below | verified live 2026-10-05 |
| Argon2 command without Rust | deploy/standalone/README.md |
cited; the API accepts any Argon2id PHC string (code: PasswordHash, algorithm check) |
The Docker images do not contain tayga-devtools |
code: the COPY line in docker/Dockerfile lists six binaries |
verified |
| Two replicas split 6 and 6; scale back to one; summed rate within 2 % | Plan 8 and Plan 9 rows below | verified live 2026-10-06 |
| ClickHouse TTLs | migrations 0001 to 0012 |
verified |
Topic retention 24 h, -1, validation, never altered |
code: default_retention_ms, validate, create_topics; tests validate_checks_retention |
verified |
| The 2026-10-06 disk incident and recovery (99 %, 28.8 → 9.1 GB, 92 % → 70 %) | Plan 9 rows below | verified live / cited where marked |
| Rows older than the TTL dropped at insert | ttl_only_drop_parts and the TTL clauses on the raw tables in crates/tayga-store/migrations |
inferred from code; not reproduced |
| Remine numbers (8.6 M logs, 501 s debug dry run, 211 s release run, 415 → 402) | Plan 7b rows below | verified live 2026-10-06 |
| Metric names and help texts | live /metrics of every service, 2026-10-07; code for the labelled families |
verified |
| 12 Pipeline charts | code: ui/src/features/pipeline/model.ts; Sub-project 4 rows |
verified |
| Grafana dashboards, read-only user | Plan 4 and Plan 9 rows below | verified live |
| Performance tuning numbers | docs/perf/sp4-performance.md |
verified (measured) |
| Upgrade notes (migration order, switch-over gap 0, republish nothing) | code: Compose depends_on; Plan 8 and Plan 9 rows below |
verified |
| Every API route, parameter, default and limit | code: routes.rs, routes_v2.rs, auth.rs, params.rs |
verified |
| Every API example and error body | live curl against the demo stack, 2026-10-07 13:30–14:10 UTC |
verified |
pipeline/series points in Unix milliseconds |
code: SeriesView doc; live response |
verified |
Performance and comparison
Section titled “Performance and comparison”| Claim | Evidence | Result |
|---|---|---|
| Every number on Performance | docs/perf/sp4-performance.md (criterion and live measurements, 2026-10-06 and 2026-10-07), Sub-project 4 rows below |
verified (measured once; not guarantees) |
GPU slower than parallel at every size |
same report, one criterion run on 2026-10-07 | verified |
| Web app budgets | Plan 5 rows below | measured once, 2026-10-05 |
| Every competitor cell | docs/research/competition.md, vendor sources accessed 2026-10-07, linked on the comparison |
verified at the source; can change |
| Log-pattern alerting is not unique to Tayga: Datadog Watchdog Log Anomaly Detection (at intake, new and increasing warning or error patterns, Watchdog logs monitor); New Relic NRQL and anomaly alerts on log patterns | Datadog logs/explorer/watchdog_insights and watchdog/alerts; New Relic find-unusual-logs-log-patterns; Dynatrace alerting-on-logs and lma-logs-app/patterns, all re-fetched 2026-10-07 after the final review |
corrected: an earlier comparison said no source described an alert on a new log pattern |
| Tayga: Docker Compose and Helm | deploy/standalone/compose.yaml; deploy/helm/tayga/Chart.yaml present on the docs branch |
files present; the chart’s own test is part of its release work |
Development record
Section titled “Development record”The rows below were kept in the README while Tayga was built, plan by plan, each checked on the branch of that plan with the stack running. They are reproduced here unchanged, so nothing is lost. Rows about the removed server-rendered pages are kept as history and marked superseded; where a later plan changed a behaviour, the row says so.
Plan 3 (API, web UI, metrics, e2e) and earlier
Section titled “Plan 3 (API, web UI, metrics, e2e) and earlier”Checked 2026-10-03 on branch feat/plan-3-api-ui-e2e, with the stack running.
| Claim | How verified | Result |
|---|---|---|
| Port 8090 = tayga-api | ports in deploy/compose.tayga.yaml; default http_addr in crates/tayga-api/src/main.rs; live curl localhost:8090/healthz returned ok |
verified |
| Port 3001 = Grafana | ports in deploy/compose.tayga.yaml; live curl localhost:3001/login returned HTTP 200 |
verified |
| Port 19090 = Prometheus | ports in deploy/compose.tayga.yaml; live curl localhost:19090/-/ready returned “Prometheus Server is Ready.” |
verified |
| Port 18123 = ClickHouse HTTP | ports in deploy/compose.infra.yaml; live curl localhost:18123/ping returned Ok. |
verified |
| Port 19092 = Redpanda Kafka API | ports in deploy/compose.infra.yaml; docker ps shows 127.0.0.1:19092->19092/tcp |
verified (port mapping only; no Kafka client connection made) |
| Tayga’s ports bind 127.0.0.1; the demo publishes 8080, 9090, 10000 and ephemeral service ports on all interfaces | lsof -nP -iTCP -sTCP:LISTEN showed 127.0.0.1:8090, :3001, :19090, :19092, :18123 and *:8080, *:9090, *:10000, *:574xx, *:627xx, *:648xx; docker ps mapped the * listeners to containers of compose project opentelemetry-demo (frontend-proxy, prometheus, otel-collector, flagd and the demo services) |
verified live 2026-10-03 |
| Port 8080 = demo frontend proxy | docker ps shows frontend-proxy on 8080; curl localhost:8080/ returned HTTP 200; the compose definition is in the submodule, not read |
verified live, definition not read |
Jaeger UI at /jaeger/ui on 8080 |
default jaeger_url in crates/tayga-api/src/main.rs and TAYGA__JAEGER_URL in compose |
verified in config only, URL not fetched |
Grafana anonymous Viewer, admin password admin |
GF_AUTH_ANONYMOUS_*, GF_SECURITY_ADMIN_PASSWORD in deploy/compose.tayga.yaml |
verified in config; login not tried; since plan 5 Grafana runs only after make up-extras; plan 9 read the dashboards through Grafana’s HTTP API without credentials |
| API and UI routes and their query parameters | crates/tayga-api/src/routes.rs, ui.rs, params.rs |
verified in code; superseded for the HTML routes (ui.rs is removed, see the Plan 5 rows) |
since range 1s-7d, defaults 1h / 24h / 1h |
parse_since, group_filter, group, service_map |
verified in code; live since=8d returned 400, since=1h&kind=error returned 200 |
kind accepts only error or slow |
group_filter in params.rs |
verified in code |
/metrics on tayga-api |
tayga_common::metrics::router merged in main.rs; live curl localhost:8090/metrics returned Prometheus text |
verified |
| Unknown route returns plain 404 | live curl localhost:8090/nope returned 404 (body not inspected); no fallback in the routers |
verified status; “plain” inferred from code; superseded: unknown paths now serve the app (see the Plan 5 rows) |
/ UI returns 200 |
live curl localhost:8090/ |
verified (still 200, now the web app) |
make up/down/ps/logs/flag/flags-reset/it/e2e/verify-raw/capture/infra-down |
Makefile |
verified in Makefile; only up-state commands were observed, make it, make e2e, make down, make infra-up and flag changes were not run |
make it conflicts with the full stack |
both compose files publish 19092 and 18123 on the host (compose.infra.yaml) |
inferred from port definitions, not run |
| Demo pinned to 3.1.0 | .gitmodules, git submodule status shows (3.1.0); DEMO_VERSION in Makefile |
verified |
| Architecture diagram | services in deploy/compose.tayga.yaml; topic tayga.signals is the default topic in crates/tayga-kafka/src/lib.rs |
topology matches compose services; topic names not checked against running Redpanda |
| Crate list | crates/*/Cargo.toml |
verified |
Plan 4 (log templates and alerts)
Section titled “Plan 4 (log templates and alerts)”Checked 2026-10-04 on branch feat/plan-4-log-templates, unless the row gives another date.
| Claim | How verified | Result |
|---|---|---|
| Port 14318 = tayga-ingest OTLP/HTTP | "127.0.0.1:14318:4318" in deploy/compose.tayga.yaml; emit-log default endpoint in crates/tayga-devtools/src/main.rs; the e2e probe sends through it |
verified in config; the e2e probe’s 45 s alert (see below) is evidence it works |
tayga-logminer runs as one replica, metrics on 9100, scraped by Prometheus |
service in deploy/compose.tayga.yaml; default metrics_addr in crates/tayga-logminer/src/main.rs; deploy/prometheus/prometheus.yml; live docker ps shows tayga-logminer Up; /api/v1/targets showed 6 jobs, all up (writer, ingest, logminer, assembler, api, redpanda) |
verified; superseded for “one replica” and the container name by plan 8 (see the plan 8 rows) |
| Rule defaults (5 min window, 60 min baseline, factor 5, min count 10, warmup 15 min measured back from the template’s first log, active 10 min, new = first seen after the previous tick’s data clock minus 60 s, start watermark 10 min before the data clock, 65 min = baseline + window) | DetectConfig::default, min_age, is_new, initial_watermark and NEW_TEMPLATE_MARGIN_NS in crates/tayga-drain/src/detect.rs, with exact-boundary unit tests; logminer settings take the same values (LogminerSettings::default and its test) |
verified in code; the e2e scenarios exercise one new-template and one spike case, not every threshold |
| Drain defaults (similarity 0.5, depth 4, 100 children, 5,000 clusters per service, 64 tokens) | DrainConfig::default in drain.rs, MAX_TOKENS in preprocess.rs |
verified in code |
Flush at 5,000 logs or 1 s, detect every 60 s, topic tayga.alerts |
LogminerSettings::default |
verified in code |
| TTLs 3 d (hits) / 30 d (templates) / 7 d (alerts) | TTL lines in crates/tayga-store/migrations/0004* |
verified in code |
New routes and their defaults (since 24h / 1h / 24h, 200 alert and template limit, q at most 200 chars) |
routes.rs, ui.rs, params.rs, repo.rs (limit 200 at log_alerts and TEMPLATES_IN_WINDOW) |
verified in code |
/alerts and /templates return 200; /api/v1/log-alerts?since=24h |
live curl: 200, 200; 8 alerts in the last 24 h |
verified live 2026-10-04; superseded: /alerts and /templates are no longer redirected and show the app’s not-found page |
| About 60-120 templates | live log_templates FINAL: 293 rows in total (all ever mined, 30-day TTL), 72 with last_seen in the last hour; /api/v1/log-templates?since=1h returned 72 |
verified live; the 60-120 range is the plan’s estimate, the live hourly count (72) is inside it |
Golden Drain test: 64 templates on the 5,000-line sample, frontend-proxy 5, bound is 120 and 10 |
cargo test -p tayga-drain --test '*' -- --nocapture printed templates: 64 {... "frontend-proxy": 5 ...}, 3 passed |
verified |
| Restoring the first half of the golden sample and mining the rest gives every line the same template id as one pass | restore_mid_corpus_matches_a_single_pass in crates/tayga-drain/tests/golden.rs. It fails (116 of 5,000 lines differ) when the restore is skipped. It still passes when clusters are restored in reverse order, so the sample does not exercise leaf-order ties |
verified 2026-10-04 |
Grafana tayga-logs (“Tayga · Logs”) has 5 panels: Log alerts by kind, Recent alerts, New templates per service, Top templates, Logs mined/s |
/api/dashboards/uid/tayga-logs, 2026-10-04 |
verified live (rendering in a browser not checked) |
tayga-pipeline has 13 panels, 5 of them logminer panels (logs mined/s, templates, cluster cap hits, detect p99, data lag) |
/api/dashboards/uid/tayga-pipeline returned 13 panels, the last “Logminer data lag”, 2026-10-04 |
verified live (rendering in a browser not checked) |
tayga_logminer_data_lag_seconds is exported and small while the logminer keeps up |
Prometheus query after make up returned 0.38 s, and 0.44 s about 40 minutes later |
verified live 2026-10-04 |
| A template that first appears while the logminer is stopped is reported after it restarts | docker stop tayga-logminer at 14:09:06; probe 1 emitted at 14:09:06; probe 2 emitted at 14:21:19, after 12 min; docker start at 14:21:19. Both new alerts were written on the first detection tick, at 14:22:19, 60 s after the start. Probe 1 was 13 min old by then, so the old wall-clock 10-minute rule would have dropped it |
single run, verified live 2026-10-04 |
Settings come from TAYGA__SECTION__KEY environment variables |
load_settings in crates/tayga-common/src/lib.rs:20-30 |
verified in code |
new_template_recent_min (10 min: the start watermark offset, the lag-warning threshold and the minimum example window) is not configurable |
LogminerSettings has no such key; the value comes from DetectConfig::default |
verified in code |
| e2e: 9 tests, 2 of them new; timeouts 180 s default, 600 s shipping, 600 s spike, 180 s new template | cargo test -p tayga-e2e -- --ignored --list listed 9; crates/tayga-e2e/src/lib.rs |
verified; superseded for shipping (300 s since plan 9, with its own orders) and ad (300 s) |
Latest make e2e run, 2026-10-04, on commit 7c0d35a: 8 of 9 passed. shipping_slowdown_produces_slow_story_blaming_shipping found no shipping slow story within 600 s. Run alone right after, on the same commit, it passed in 40 s. Alert times: log spike 200 s; new-template probe 45 s, after a 556 s first-run warmup wait for the tayga-e2e-probe service |
one make e2e run plus one single-test re-run |
single run, not a latency guarantee; the shipping failure fits the rarity of international orders (see the later make e2e rows) |
| Story scenario times in that run: ad 85 s, payment 95 s, unreachable 65 s, catalog 25 s, shipping failed (40 s in the re-run) | same run | single run; earlier runs differed |
| Performance, scale, or latency claims | none made beyond the single-run timings above | n/a |
Plan 5 (web app)
Section titled “Plan 5 (web app)”Rows below checked 2026-10-05 on branch feat/plan-5-ui at 870bfaa, against the stack from make up (5 tayga containers running; Grafana and Prometheus not running).
| Claim | How verified | Result |
|---|---|---|
App served on 8090: /, /map, /logs, /logs/alerts, /pipeline, /traces return 200 text/html; an unknown path (/nope) also returns 200 text/html |
live curl -D - |
verified |
/api/v1/nope and /assets/nope.js return 404 application/json; /metrics returns 200 |
live curl |
verified |
Old-URL redirects (/groups/…, /service-map, /alerts, /templates) |
removed in plan 6 (3dffb27); old_urls_are_plain_client_routes in crates/tayga-api/src/spa.rs |
superseded: no redirects; old URLs are client routes and show the not-found page |
GET /api/v1/config returns {"jaeger_url":"http://localhost:8080/jaeger/ui","grafana_url":null} on a plain make up |
live curl |
verified at the time; superseded: the response now also has auth_enabled and infra_services, {"jaeger_url":"http://localhost:8080/jaeger/ui","grafana_url":null,"auth_enabled":false,"infra_services":["flagd"]} (see the rows below and the HTTP API) |
New API routes overview, stories/series, traces/search, search?q=, pipeline/series return 200; services returns a list of names; pipeline/lag returns three groups (writer, assembler, logminer) |
live curl (pipeline/series?metric=up&kind=gauge&job=tayga-api&since=15m; services/{name} not called) |
verified live |
Route list, parameters and limits (limit 1 to 500, default 100; touched, errors flags; kind required for pipeline/series) |
crates/tayga-api/src/routes_v2.rs, params.rs |
verified in code |
Immutable cache on /assets/*, no-cache on index.html |
IMMUTABLE and NO_CACHE in spa.rs; live curl -D - on 2026-10-04: cache-control: public, max-age=31536000, immutable on the asset, no-cache on / |
verified in code and live on 2026-10-04; not re-fetched today |
Pages, paths, shortcuts (g then s t m l p, ?, Cmd/Ctrl+K), theme cycle light, dark, system, time ranges 15m/1h/24h/7d, live refresh 10 s paused while hidden, palette contents |
ui/src/router.tsx, components/shell/{Shortcuts,CommandPalette,ThemeSwitch,TimeRange,LiveToggle}.tsx, app/search.ts, theme/theme.ts |
verified in code; not clicked through in a browser today (the Playwright specs ui/e2e/pages.spec.ts, theme.spec.ts and palette.spec.ts cover pages, theme switch and palette) |
Recorder: every 15 s, 7-day TTL, targets in deploy/tayga-api.toml via TAYGA_CONFIG, not settable by TAYGA__ env |
default_record_secs in crates/tayga-api/src/main.rs; TTL ... INTERVAL 7 DAY in 0005_metric_samples.sql; live 2026-10-04: config 0.15.27 rejected the env form with invalid type: map, expected a sequence |
verified in code; the env failure was seen once on 2026-10-04, not re-run |
Grafana and Prometheus are the compose profile extras; make up-extras starts them and sets the Grafana link; make down removes them |
Makefile, deploy/compose.tayga.yaml, deploy/compose.extras.yaml; live 2026-10-04: up-extras gave grafana_url set, Prometheus ready, 4 dashboards provisioned; a second make up left them running |
verified in files and live on 2026-10-04; make up-extras not run today |
Node 24 or newer only for UI development; Docker builds with node:24; make ui-dev, make ui-e2e |
engines in ui/package.json; first stage of docker/Dockerfile; Makefile; node --version here prints v24.18.0 |
verified |
Without ui/dist, tayga-api compiles and serves a placeholder |
spa.rs module comment, PLACEHOLDER, allow_missing |
verified in code; a build without ui/dist not run today |
Dev server proxies to 8090 by default, TAYGA_API overrides |
ui/vite.config.ts |
verified in code |
| UI unit tests: 371 pass | npm --prefix ui test run today: Tests 371 passed (371) |
verified |
| Playwright suite: 139 tests in 11 files across 4 projects (dark, light, reduced motion in each theme) plus the perf project | Full run on 2026-10-05 against the 8090 app rebuilt from the final branch code (make up, image built 08:50 UTC): 139 passed, 12 skipped, 0 failed in 1.3 min. The 12 skips are the motion specs, which run only in the reduced-motion projects |
one full run |
Budgets (as asserted in ui/e2e/perf.spec.ts), final run on 2026-10-05, unthrottled on the development machine against the live stack, medians of 5 cold-cache contexts: initial JS 185.2 KB gzip (limit 350 KB); JS fetched by a cold home load 231.1 KB (350 KB); home first render with data 146 ms (1000 ms); waterfall of the largest live trace (105 spans) 25 ms and of 5,000 synthetic spans 32 ms (200 ms); map layout 143 ms live and 181 ms for 60 synthetic nodes (300 ms); live refresh gaps 10071 and 10048 ms; 0 requests in 13 s while the tab is hidden |
Playwright perf project output of that run |
measured once on one machine, not a guarantee |
Last full make e2e, 2026-10-05, after the scenario fixes in 3d0d8a5: 8 of 9 passed. shipping_slowdown_produces_slow_story_blaming_shipping stopped at its baseline pre-check: two shipping runs in the previous hour lifted checkout p99 to 1.26 s, and each checkout endpoint had 35 to 37 traces in the window, below the 50 the detector needs. Run alone about 20 minutes earlier, with a clean baseline, it passed in 160 s. Earlier timeouts in this plan’s Playwright runs came from the demo checkout service running about 100 times slower after a Docker restart; restarting checkout and load-generator fixed it |
one full run plus single-test runs | the shipping scenario needs an hour without slowdown runs and enough checkout traffic |
| Performance, scale, or latency claims | none beyond the cited budgets and single runs above | n/a |
Plan 6 (login, infrastructure toggle, no redirects, agent noise)
Section titled “Plan 6 (login, infrastructure toggle, no redirects, agent noise)”Rows below checked 2026-10-05 on branch feat/plan-6-owner-decisions at a6cf80a (plus this docs commit). Auth rows were run against a host tayga-api debug build on 127.0.0.1:18090 (ClickHouse :18123, Kafka :19092), stopped afterwards.
| Claim | How verified | Result |
|---|---|---|
[auth] keys enabled, username, password_hash, session_ttl (default 12h), session_key, secure_cookie, and TAYGA__AUTH__* env forms |
AuthSettings and its Default in crates/tayga-api/src/auth.rs; live: the host API started with only TAYGA__AUTH__ENABLED, TAYGA__AUTH__USERNAME, TAYGA__AUTH__PASSWORD_HASH set accepted logins and returned Max-Age=43200 (12 h) |
verified |
session_ttl is capped at 365d; username may not contain | or :; session_key must be at least 32 bytes base64 |
live startup errors with the env forms: auth.session_ttl must be <n>s, <n>m, <n>h or <n>d, above 0 and at most 365d for 366d; auth.username must not contain '|' or ':' for a|b and a:b; auth.session_key must decode to at least 32 bytes for 3 bytes. With 365d, a 32-byte key and secure_cookie=true the login returned Max-Age=31536000 and Secure, and the log had no “session_key is unset” warning |
verified live |
Without session_key a random key is used and a restart logs everyone out |
startup log of the host API: WARN auth.session_key is unset: a random key is used, so a restart signs everyone out; key generation in Auth::from_settings |
verified live (log line); that old cookies stop working after a restart is by construction (the key changes), not re-tried today |
| The session is not sliding | the cookie’s Max-Age is set only at login (crates/tayga-api/src/auth.rs); no route re-issues it |
verified in code; not tested by waiting out a session |
hash-password prints an Argon2id hash from piped stdin; asks twice on a TTY |
live: printf 'secret' | tayga-devtools hash-password printed a string starting $argon2id$v=19$m=19456,t=2,p=1, and the API accepted it; read_password and confirmed in crates/tayga-devtools/src/password.rs; unit test hash_of_known_password_verifies |
piped form verified live; the two-prompt TTY path verified in code only (no TTY here) |
Login returns 204 with a tayga_session cookie HttpOnly; SameSite=Strict; Path=/; protected routes need it; Basic works for scripts |
live on :18090: GET /api/v1/story-groups without credentials 401; POST /api/v1/auth/login with a JSON body 204 with Set-Cookie: tayga_session=…; HttpOnly; SameSite=Strict; Path=/; Max-Age=43200; curl -u admin:secret …/story-groups?since=15m 200 |
verified live |
Open routes: /healthz, /metrics, GET /api/v1/config, GET /api/v1/auth/me (401 without a session, so reachable), POST login and logout, and the app’s files |
live on :18090 with no credentials: /healthz 200, /metrics 200, /api/v1/config 200 {"jaeger_url":null,"grafana_url":null,"auth_enabled":true,"infra_services":["flagd"]}, / 200, /api/v1/auth/me 401; OPEN_API in auth.rs; tests open_routes_match_method_and_exact_path and open_routes_stay_open (cargo test auth::: 28 passed) |
verified live and in tests |
| Logout needs a JSON content type and no session | live: POST /api/v1/auth/logout without a content type 415; with application/json and no cookie 204; logout_requires_json_content_type and logout_needs_no_session in auth.rs |
verified live |
| Limiter: 5 failed attempts per client IP per 5 minutes, the 6th is refused even with the right password | live: 5 wrong logins all 401, then the right password 429 with retry-after: 299 and {"error":"too many attempts"}; WINDOW and MAX_FAILURES in auth.rs |
verified live (IPv4); IPv6 /64 keying verified in the unit test ipv6_is_limited_per_64_prefix, not live |
X-Forwarded-For is ignored, so behind a proxy all users share one limit |
forwarded_for_does_not_dodge_the_limiter in auth.rs; live run on 2026-10-05: 5 wrong logins with a different X-Forwarded-For each, the 6th was 429 |
verified in a test and live on 2026-10-05; not re-run today |
| Parallel attempts are counted before the password check, so more than 5 in flight can get 429 | the doc of Limiter in crates/tayga-api/src/auth.rs (“An attempt is recorded before the check runs”) and Limiter::attempt |
verified in code; not measured |
Login page, redirect to /login?next=…, wrong-password alert, sign-in, reload stays signed in, sign-out, an unsafe next lands on / |
ui/e2e/auth.spec.ts run today against the Vite dev server (5174) proxied to the auth API (18090): TAYGA_E2E_AUTH_USER=admin TAYGA_E2E_AUTH_PASS=… TAYGA_UI_URL=http://localhost:5174 npx playwright test auth.spec.ts --project=dark --project=light --no-deps |
verified: 8 passed (4 tests in dark and light) |
HTTPS goes through a reverse proxy with secure_cookie |
secure_cookie adds ; Secure (live: login with TAYGA__AUTH__SECURE_COOKIE=true returned a cookie ending Secure); tayga-api binds plain HTTP (http_addr, no TLS code in main.rs) |
cookie flag verified live; no reverse proxy was set up or tried |
[map] infra_services defaults to ["flagd"], is in /api/v1/config, is settable in file config only |
live: /api/v1/config on the rebuilt stack (make up) returned "infra_services":["flagd"]; with a TOML file [map] infra_services=["flagd","otel-collector"] (TAYGA_CONFIG) it returned both; with TAYGA__MAP__INFRA_SERVICES=flagd,otel-collector the API failed at startup with invalid type: string "flagd,otel-collector", expected a sequence for key map.infra_services`` |
verified live |
“Show infrastructure” toggle stored in the URL as infra=true; flagd hidden by default and drawn with it; ?service=flagd opens the drawer while hidden |
ui/e2e/map.spec.ts tests infrastructure (flagd) is hidden by default and drawn with infra=true and /map?service=flagd opens flagd even while infrastructure is hidden, in the full Playwright run below; MapSearch.infra in ui/src/app/search.ts |
verified |
| Callers of hidden infrastructure keep a “+N infra” badge; the header’s degraded count still includes infrastructure | hideInfra in ui/src/features/map/model.ts; unit tests in model.test.ts and map.test.tsx (the header count is 4 against the map’s 3 in the fixture); ui/e2e/map.spec.ts asserts a visible [data-infra-badge] while infrastructure is hidden |
verified in unit tests and Playwright |
| Old-URL redirects removed | old_urls_are_plain_client_routes in crates/tayga-api/src/spa.rs; commit 3dffb27 |
verified in a test (see also the superseded rows above) |
The load generator no longer calls the agent service: no user_ask_agent spans after the deploy while other spans flow |
docker exec load-generator env shows LOCUST_LOCUSTFILE=/usr/src/app/tayga_locustfile.py. ClickHouse tayga.spans, window 13:17:05-13:23:05 UTC on 2026-10-05, after the Docker disk was freed: user_ask_agent 0, other load-generator spans flowing (GET 184, POST 99, user_browse_product 65, checkout 16). Rechecked at 13:31:49 UTC, last 20 minutes: user_ask_agent 0, 521 spans named user_%, 105,313 spans in all. The last user_ask_agent span was 2026-10-05 11:41:59, before the 11:48 deploy |
verified live; the 20-minute window is short and the demo’s other tasks are random |
| Demo memory limits raised; all services below 65% after the change | docker compose … config --format json shows the new limits; docker stats 2026-10-06 after make up: checkout 23/64 MiB, product-catalog 29/64, ad 277/512, fraud-detection 253/512, accounting 190/320, kafka 596/1024, opensearch 997/1536, grafana 148/256, load-generator 660/2048; no unhealthy containers. Before: checkout 19/20 (95%), ad 272/300, load-generator 1.27/1.47 GiB |
verified live; one snapshot |
| The agent service is not part of Tayga’s stack | the agent service exists only in the demo’s compose.agent.yaml; the Makefile’s COMPOSE lists compose.yaml, compose.full.yaml, compose.observability.yaml and Tayga’s two files |
verified in files |
make up on this branch builds an API with the new config fields |
make up finished in 24 s; curl localhost:8090/api/v1/config returned auth_enabled:false and infra_services:["flagd"] |
verified |
| Rust unit tests: 306 pass, 29 ignored | cargo test --workspace run today (sum of the test result lines) |
verified |
| UI unit tests: 463 pass | npm --prefix ui test run today: Tests 463 passed (463), 37 files |
verified |
| Playwright on the 8090 app with auth disabled: 131 passed, 28 skipped, 0 failed | npx playwright test from ui/ run today after make up: 131 passed (1.6m). The skips: 16 for auth.spec.ts (needs the auth env) and 12 motion specs (reduced-motion projects only) |
one full run |
make e2e on 2026-10-05 after make up: 8 of 9 passed in 497 s. Story times: ad 75 s, payment 55 s, unreachable 70 s, catalog 25 s; log spike 211 s; new-template probe 60 s after 0 s warmup (the probe service was already warm). shipping_slowdown_produces_slow_story_blaming_shipping stopped at its baseline pre-check: checkout baselines cannot flag a 5 s trace as slow with an empty endpoint table |
make e2e output; make flags-reset run after |
single run, not a latency guarantee. The test itself marked ad (75 s) and unreachable (70 s) as over its 60 s target; both passed |
Plan 7a (detection correctness)
Section titled “Plan 7a (detection correctness)”Rows below checked 2026-10-05 on branch feat/plan-7a-detection at 42d9dbb (plus this docs commit), against the stack redeployed by make up at 16:10:20 UTC (28 s; migrations 0006 to 0008 applied). Docker disk 85% before the deploy.
| Claim | How verified | Result |
|---|---|---|
Migrations 0006, 0007, 0008 applied; the three logminer_state keys exist |
live SELECT key, value FROM logminer_state FINAL: masking_epoch_start_ns 1791216620578814135 (16:10:20.58 UTC, the logminer’s start), masking_version 2, new_template_watermark_ns 1791217400436115384 (16:23:20, advancing per pass); show tables lists logminer_state, log_template_minutes, log_template_minutes_mv; describe log_alerts shows baseline_day, baseline_week |
verified live; superseded for masking_version (now 3, see the v3 rows below) |
| An upgrade over existing templates starts a masking epoch | logminer start log: restored: 384, masking_version: 2, epoch_start = 1791216620578814135; the DB had templates and no stored version |
verified live (v2 deploy); superseded: the v3 deploy started a new epoch at 18:33:56, see below |
| No new-template alerts after the deploy during the epoch warmup | live SELECT count() FROM log_alerts FINAL WHERE kind='new' AND started_at BETWEEN '16:10:20' AND '16:25:20': 0 (all services), 0 for frontend-proxy. Alerts in that window: one spike (otelcol-contrib, “Exporting failed. Will retry the request after interval.”). The first new alert after the deploy came at 17:12 (load-generator, a currency-task timeout), outside the window |
verified live (one window, the v2 deploy). After the v3 epoch a rarer combination alerted once past the warmup, see the 18:49:39 row below |
log_template_minutes fills |
live SELECT count(), uniqExact(template_id) FROM log_template_minutes: 847 rows at 16:24, 6091 rows over 68 templates at 17:54 |
verified live |
Existing frontend-proxy templates do not gain literal status codes on upgrade |
live: 0 frontend-proxy templates matching a 3-digit 1xx-5xx token and none with ' 200 ', 6 templates in total, none first seen after the deploy, while the stored log bodies contain HTTP/1.1" 200; templates still read "GET <*> <*> <*> ... |
superseded (v2 behaviour). Since ac8e775 a kept code never matches <*>, so access-log lines with a code start new templates even on existing services; no template wipe or re-mining is needed. See the v3 rows below |
| A new service’s access-log template keeps the status code | live: emit-log --service tayga-status-probe with an Envoy-style body containing HTTP/1.1" 503 UF produced the template <*> "GET /api/cart <*> 503 UF <*> <*> <*> - "-" "probe" in log_templates |
verified live (one body, one service; the probe service stays in the table until its 30-day TTL) |
Status rule details (100 to 599, only after an HTTP/x token, 600 and non-HTTP numbers masked, per-token equals whole-string masking) |
is_status, is_http_version in crates/tayga-drain/src/preprocess.rs; unit tests keeps_status_after_http_version, other_numbers_stay_masked, per_token_masking_equals_whole_string_masking |
verified in code and unit tests |
| New metrics are exported | live: assembler /metrics showed tayga_assembler_baseline_excluded_traces 2, tayga_assembler_baseline_carried_endpoints 0, tayga_assembler_baseline_endpoints 56; logminer /metrics showed tayga_logminer_spike_skipped_total{reason="coverage"} 0, tayga_logminer_seasonal_failures_total 0, tayga_logminer_state_save_failures_total 0 (scraped with curlimages/curl on the opentelemetry-demo network, because the Tayga images have no wget or curl) |
verified live; the counters are 0, so no skip, failure or carry was observed live |
| Coverage rule, per-template coverage, young-template baseline, proportional shortening, the 50% gate | template_coverage in crates/tayga-drain/src/detect.rs and spike_baseline; unit tests thin_coverage_is_skipped, coverage_shrinks_the_baseline_windows, young_templates_can_spike_after_ten_minutes, a_young_template_spanning_an_outage_is_not_judged, coverage_before_the_template_existed_does_not_count; store ITs in crates/tayga-store/tests/store_it.rs (18 passed against live ClickHouse on 2026-10-05) |
verified in code and tests; no live coverage skip occurred (counter 0) |
Seasonal mode: 1-day and 7-day comparators, optional baseline_day/baseline_week, flat fallback on lookup error, MV dedup under replay |
store ITs against live ClickHouse on 2026-10-05 (20 passed), including minutes_mv_does_not_double_count_replayed_hits; baseline_mode_setting_is_validated in crates/tayga-logminer/src/main.rs |
verified in unit and store tests; not run live (at the time the stack had under a day of log_template_minutes; since migration 0009 it holds minutes back to 2026-10-02 17:37; the deployed mode is flat) |
| Slow baselines exclude above-cap traces; endpoints are carried for at most 2 windows; no baseline from no traces | crates/tayga-store/src/store.rs, crates/tayga-assembler/src/baselines.rs; unit tests and store ITs (2026-10-05), including one_outlier_does_not_raise_p99_with_previous_cap and bootstrap_cap_is_ten_times_p50) |
verified in code and tests; live: 2 excluded traces and 0 carried endpoints at one scrape, no sustained slowdown observed |
| Carry state is in memory only; a slowdown longer than 2 windows stops being flagged | baselines.rs (no persistence) |
code only, not run live |
| Rust unit tests: 348 pass, 36 ignored | cargo test --workspace run today (sum of the test result lines) |
superseded: 365 pass, 37 ignored after the final-review fixes, see below |
make e2e on 2026-10-05 after the 7a deploy (about 2 h after make up): 7 of 9 passed in 647 s. Passed: log spike 175 s, new-template probe 60 s after 0 s warmup, payment 125 s, unreachable 80 s, catalog 25 s, raw span counts, service map. Failed: ad_failure_blames_ad (one story in 180 s; the test needs 3) and shipping_slowdown_produces_slow_story_blaming_shipping (stopped at its pre-check with an empty table) |
make e2e output; make flags-reset run after |
single run |
| The ad failure was not a regression; re-run alone it passed in 80 s | cargo test -p tayga-e2e -- --ignored ad_failure_blames_ad: adFailure: first matching story after 80s, 1 passed |
single re-run; the first failure’s cause is not established (the flag affects about one request in ten, so a slow draw is plausible) |
| The shipping pre-check stop is environmental | live trace_summaries for the last 60 min: user_checkout_single 73 of 73 and user_checkout_multi 59 of 59 traces had is_error = 1; the 10-minute counts show 100% checkout errors every interval since 13:10 UTC, hours before the deploy; docker logs checkout shows panic: runtime error: invalid memory address or nil pointer dereference. The pre-check needs non-error checkout traces, so no rows came back |
error counts verified live. Corrected: the only panic in docker logs -t checkout is stamped 2026-10-04 18:00:22 UTC, a day earlier, so it is not the cause; the evidence points to the demo’s Kafka (slow publish orders, recovery after its restart), see the checkout rows below |
Plan 7a final-review fixes (v3 masking)
Section titled “Plan 7a final-review fixes (v3 masking)”Rows below checked 2026-10-05 between 19:12 and 19:26 UTC on feat/plan-7a-detection with the fix commits after ac8e775, against the stack redeployed by make up at 19:16:29 to 19:16:55 UTC.
| Claim | How verified | Result |
|---|---|---|
masking_version 3, masking epoch at 2026-10-05 18:33:56 UTC |
live SELECT key, value FROM logminer_state FINAL: masking_epoch_start_ns 1791225236728986679 (fromUnixTimestamp64Nano: 18:33:56.729), masking_version 3; the 19:16 restart logged restored: 399, epoch_start: 1791225236728986679, masking_version: 3 (unchanged version, so no new epoch) |
verified live |
7 new frontend-proxy templates carry a literal status, next to the 6 older <*> ones |
live log_templates FINAL for frontend-proxy: 13 templates, 7 with a [1-5][0-9][0-9] token, codes 200 (4), 308, 503 and 504, first seen (log time) from 18:33:53.51 to 18:49:39.43; the other 6 were first seen 2026-10-02 05:37 to 2026-10-03 17:59. The first four are stamped up to 3.2 s before the epoch because lines logged before the logminer’s start were mined after it |
verified live |
The 18:49:39 new alert for <*> "GET <*> <*> 503 UC upstream_reset_before_response_started{connection_termination} ... was a false positive |
log_alerts: kind='new', frontend-proxy, started_at 18:49:39.434. logs holds 14 frontend-proxy bodies containing 503 UC before the epoch (2026-10-02 16:16 to 2026-10-05 17:10; 13 of them still have hits, 6 in template 6470045815083748258 and 7 in 11039615203255878215, both with <*> at the status); 1 after it. The template is new only because the kept 503 no longer matches <*> |
verified live. Fixed by the pre-epoch match (final review I1): replaying the rule over the live log_templates (status read as <*>, same length and first two tokens, similarity ≥ 0.5) matches 11039615203255878215 with similarity 1.0, so the alert would now be suppressed. Unit tests a_status_split_out_of_a_pre_epoch_wildcard_template_would_have_matched, new_alerts_for_status_splits_of_pre_epoch_templates_are_suppressed |
| A genuinely new status shape still alerts; templates without a kept code are judged as before | unit tests a_genuinely_new_shape_with_a_status_would_not_have_matched, templates_without_a_protected_token_are_never_suppressed, pre_epoch_match_survives_a_restore_and_spares_non_http_templates |
verified in unit tests; not observed live |
tayga_logminer_new_suppressed_total{reason="pre_epoch_match"} is exported |
logminer /metrics after the 19:16 deploy (scraped with curlimages/curl on the opentelemetry-demo network): tayga_logminer_new_suppressed_total{reason="pre_epoch_match"} 0 |
verified live; 0, so no suppression observed live yet |
Migration 0009 backfilled log_template_minutes without double counting |
live: schema_migrations version 9 applied 19:16:52. Before the deploy the earliest minute was 2026-10-05 16:10:00 (the view’s start, 187 minutes); after it 2026-10-02 17:37:00 (4,302 minutes). Sums of uniqExactMerge(hits) per (template, minute) equal sums of uniqExact(log_id) from log_template_hits: before 16:10, 8,433,847 = 8,433,847; the boundary minute 16:10, 2,243 = 2,243; 16:11 to 19:00, 359,986 = 359,986 |
verified live; store IT backfill_fills_minutes_before_the_view_without_double_counting |
Duration caps come only from trusted previous baselines; after a carry expires with 0 < kept < 50 the next refresh bootstraps and adopts the new level |
caps_from in crates/tayga-assembler/src/baselines.rs; unit tests caps_skip_untrusted_baselines, after_a_carry_expires_with_few_kept_the_next_refresh_adopts_the_new_level |
verified in unit tests (the 10 x p50 bootstrap itself is the store IT bootstrap_cap_is_ten_times_p50); not run live |
| An empty or failed consumer assignment holds the new-template clock at the watermark and keeps per-partition state | assigned_partitions, pass_clock in crates/tayga-logminer/src/main.rs; unit test an_empty_or_failed_assignment_holds_at_the_watermark_and_keeps_partition_state |
verified in unit tests; no live rebalance tried |
| Rust unit tests: 365 pass, 37 ignored; store ITs: 21 pass | cargo test --workspace (sum of the test result lines); TAYGA_IT_CLICKHOUSE=http://localhost:18123 cargo test -p tayga-store --test store_it -- --ignored |
verified |
Demo checkout: every user_checkout_* trace errored from about 13:00 UTC until the maintainers restarted the demo’s Kafka and checkout |
trace_summaries FINAL, 15-minute buckets: 13:15 34/34 errors, 13:30 40/40, and every bucket from 17:15 to 18:45 100% (the table keeps 2 days; earlier, 11:00 to 12:00 had 89 of 109 errors and 12:00 to 13:00 only 2 traces). Proxy: 820 of 863 POST /api/checkout frontend-proxy bodies from 13:15 to 19:05 carry 504 (template ... "POST /api/checkout <*> 504 UT response_timeout ...). checkout’s publish orders span, 13:00 to 19:05: 825 spans, median 91.5 s, max 589.4 s. Restart: docker inspect StartedAt kafka 19:07:16 UTC, checkout 19:07:57 UTC |
verified live. The demo Kafka at 600.8 of its 620 MiB memory limit is the maintainers’ docker stats reading before the restart and cannot be re-checked now (548.2 MiB / 620 MiB at 19:19) |
| Checkout recovered after the restart | SELECT toStartOfFifteenMinutes(ts), count(), countIf(is_error=1) FROM tayga.trace_summaries FINAL WHERE endpoint_name LIKE 'user_checkout%' AND ts > now()-INTERVAL 2 HOUR GROUP BY 1 ORDER BY 1 at 19:25 UTC: 18:45 40/40 errors, 19:00 40/20, 19:15 25/0; in 5-minute buckets 19:05 9/5, then 19:10 16/0, 19:15 11/0, 19:20 14/0. publish orders after 19:08: 31 spans, median under 0.1 s |
verified live (about 15 minutes after the restart) |
Trap: span status_code is lowercase |
spans.status_code is Enum8('unset' = 0, 'ok' = 1, 'error' = 2); over checkout spans 13:15 to 19:00, countIf(status_code='error') = 86 while countIf(status_code='ERROR') = 0, with no error raised |
verified live |
Plan 7b (silence alerts, notifier, remine)
Section titled “Plan 7b (silence alerts, notifier, remine)”Rows below checked 2026-10-06 from 05:39 to 07:02 UTC on feat/plan-7b-alerting (e8621ff plus the plan’s last commits), against the stack rebuilt by make up at 05:39 UTC.
| Claim | How verified | Result |
|---|---|---|
A silence alert fires in log time, and started_at is the last hit |
make e2e-notifier: templates A and B existed 5 s after the emit; silence enabled with minutes = 2; alert fe0492738da0b683 155 s after enabling, with last_at − started_at = 159 s. In make e2e: alert f59f6d2b7fec6162 after 176 s, with 176 s. The test asserts last_at − started_at ≥ minutes |
verified live (2 runs) |
| Silence UI: switch, minutes (default 10, 1 to 1440, disabled while off), bell, badge, “Silence” filter, “silent N min” | ui/src/features/logs/SilenceCard.tsx (DEFAULT_MINUTES = 10, SILENCE_MAX = 1440, disabled={!cur.enabled}), TemplatesTable.tsx (Bell), routes/logs/alerts.tsx ({ value: 'silence', label: 'Silence' }), features/logs/model.ts (silent ${min} min) |
verified in code; the UI was not driven in a browser for these rows |
PUT silence: 415, 400, 404, auth; silence and silence_enabled on the GETs |
put_silence in crates/tayga-api/src/routes.rs; router tests put_silence_statuses and put_silence_failure_is_503; silence_put_needs_a_session_or_basic in auth.rs; live PUTs (200) from both e2e runs |
verified in code and tests; 200 path live |
Partition-clock guard: s_last counts only up to the data clock |
is_silent in crates/tayga-drain/src/detect.rs (s_last.min(clock_ns)); unit test a_held_back_clock_suppresses_silence_alerts |
verified in unit tests; no lagging partition observed live |
Silence alerts are not alerting and not badges on story logs |
kind != 'silence' in ALERTING_AT and in the per-trace alert query, crates/tayga-api/src/repo.rs |
verified in code |
A silence ends when it is no longer refreshed; inactive 10 min after last_at |
active = last_at > end − ALERT_ACTIVE_MIN (10) in repo.rs; live: after the scenario switched silence off, fe0492738da0b683 showed active: false at 06:17 UTC with last_at 05:45:20 |
verified live |
host.docker.internal resolves inside the notifier on Docker Desktop for Mac |
docker exec tayga-notifier getent hosts host.docker.internal: 192.168.65.254 |
verified live; the host-gateway mapping for Linux is not tested |
| The override replaces the config mount | docker compose … -f deploy/compose.notifier-e2e.yaml config tayga-notifier: one bind of deploy/tayga-notifier.e2e.toml at /etc/tayga/notifier.toml |
verified |
The notifier delivers a fresh alert exactly once, also after docker restart tayga-notifier |
make e2e-notifier (exit 0, 310.9 s): the mock got 1 delivery of fe0492738da0b683, 0 s after the alert showed in the API (4 requests in all, the others for other alerts). After the restart and a 150 s watch, still 1. tayga.alerts partition 2 holds the alert at offsets 760, 761 and 762 (last_at 05:43:20, 05:44:20, 05:45:20), so two re-publishes came after the restart. notifier_deliveries: one row, e2e-mock, delivered, attempts 1 |
verified live (one run) |
Webhook body fields, including last_at and summary |
the delivered body above: alert_id, kind, service, template_id, template, started_at, last_at, count 0, baseline 0.0, summary “… has been silent for 2 min in tayga-e2e-probe”, example_trace_ids [], links; content type application/json (asserted by the scenario) |
verified live |
| The default config is target-less again after the check | docker logs tayga-notifier after make e2e-notifier: delivery disabled: no targets, and tayga-notifier consuming with "targets":"[]" |
verified live |
Retries, Retry-After, permanent 4xx/3xx, backoff, max_age_secs, the 40 s stop grace, URL redaction |
classify, retry_after_delay, backoff, retry_wait in crates/tayga-notifier/src/deliver.rs; route.rs; stop_grace_period: 40s in deploy/compose.tayga.yaml; the unit tests in deliver.rs and crates/tayga-notifier/tests/notifier_it.rs |
verified in code and tests; not exercised live with a failing target |
| Notifier metrics, the Pipeline job and the lag | live 2026-10-06: /metrics served the three series; pipeline/series?metric=up&job=tayga-notifier gave 1.0; pipeline/lag listed tayga-notifier with lag 0 |
verified live on 2026-10-06, not re-checked here |
Receivers should dedup on alert_id: hard-kill, lost-final-write and 30-day TTL resends |
deliver.rs module docs; TTL toDateTime(updated) + INTERVAL 30 DAY in migration 0011_notifier_deliveries.sql |
verified in code; accepted, not reproduced |
remine --dry-run on live data |
06:07:47 to 06:16:08 UTC, debug build: 8,655,171 logs read; templates 412 before, 402 after; 31 added, 41 removed, 371 unchanged by id; otelcol-contrib 35 → 21, frontend-proxy 19 → 20, tayga-e2e-probe 12 → 15; orphaned silence settings: none; duration 500.9 s (370 s user CPU) |
verified live |
| The real run refuses while the heartbeat is fresh | after docker compose … stop tayga-logminer at 06:40:01: Error: the logminer looks alive (heartbeat 43s ago, limit 180s). Stop it first with ..., exit status 1 |
verified live |
| Real re-mine | release build, 06:42:30 to 06:46:01 UTC: 8,647,835 logs read and hits written; templates 415 before, 402 after; 27 added, 40 removed, 375 unchanged; no orphaned silence; 211.1 s (33 s user CPU); it printed “Start the logminer again now.” Before: 415 templates, 8,984,121 hits, 301,878 log_template_minutes rows. After: 402 templates, 8,647,205 hits, 259,642 minute rows |
verified live |
logminer_state after the re-mine |
masking_epoch_start_ns 1791268950154855000 (06:42:30.15, was 1791225236728986679), new_template_watermark_ns 1791269158886828000 (06:45:58.89), masking_version 3; the heartbeat stayed at 06:39:21 until the restart. The logminer restart at 06:46:09 logged restored: 402, epoch_start: 1791268950154855000, masking_version: 3 |
verified live |
No new alerts during the warmup after the re-mine |
SELECT toString(kind), count() FROM log_alerts FINAL WHERE started_at >= '2026-10-06 06:46:09' AND started_at < '2026-10-06 07:01:09' GROUP BY kind: no rows, so no alert of any kind in the 15 minutes after the restart (also none up to 07:08). log_templates had 0 templates first seen after the epoch, and the logminer was live: heartbeat 07:08:11, 51,040 hits after 06:46:09 |
verified live; weak test, because no new template appeared in the window |
make e2e: 9 of 10 pass |
06:02:36 to 06:18:24 UTC, 943.9 s, exit 2. Passed: log spike 221 s, new-template probe 60 s (warm after 0 s), payment 111 s, unreachable 50 s, catalog 25 s, raw span counts, service map, shipping 115 s (its pre-check passed), silence 176 s. Failed: ad_failure_blames_ad (groups with 2 and 1 stories in 180 s; it needs 3). Run alone afterwards it passed in 85 s. make flags-reset run after both |
verified live; the ad failure is the same intermittent miss as in plan 7a |
| Rust unit tests: 428 pass, 49 ignored | cargo test --workspace (sum of the test result lines); fmt and clippy -D warnings clean |
verified |
Migrations run before the services that need them under make up |
depends_on: tayga-migrate: condition: service_completed_successfully on writer, assembler, logminer, notifier and api in deploy/compose.tayga.yaml |
verified in config |
| Performance, scale, or latency claims | none beyond the single runs above | n/a |
Plan 8 (logminer replicas)
Section titled “Plan 8 (logminer replicas)”Checked 2026-10-06 (UTC times) on branch feat/plan-8-logminer-scale.
| Claim | How verified | Result |
|---|---|---|
tayga.logs exists with 12 partitions and the same max.message.bytes as tayga.signals |
rpk topic describe tayga.logs -p lists partitions 0 to 11; rpk topic describe -c: max.message.bytes 1048576 on both, retention.ms 604800000 (default) on both |
verified live; superseded for retention: both are 24 h since 2026-10-06 16:14 UTC (see the plan 9 rows) |
| Ingest publishes logs to both topics | ingest start log: "topic":"tayga.signals","logs_topic":"tayga.logs"; ingest /metrics: tayga_ingest_log_records_published_total{topic="tayga.signals"} 49480, {topic="tayga.logs"} 10484, and records_published_total{kind="logs"} 49480 (the signals copy only) |
verified live |
tayga.logs is keyed by service: each service is in exactly one partition |
rpk topic consume tayga.logs -o :end -f '%p %k\n' | sort | uniq -c at about 09:20: 17 services over 9 partitions (0 shipping; 3 currency, fraud-detection; 4 quote; 5 accounting, frontend-proxy; 6 checkout, product-catalog, recommendation; 7 cart, load-generator; 8 ad, email, frontend, otelcol-contrib; 9 payment; 11 kafka), none in two; partitions 1, 2 and 10 empty |
verified live (one snapshot) |
| Switch-over gap on this upgrade | uniqExact(log_id) from logs and from log_template_hits for 08:30 to 09:00 (the upgrade was 08:56): 69,315 and 69,315 |
verified live: 0 logs unmined this time; the gap depends on the old logminer’s lag at the stop |
| One replica: consumes all 12 partitions and keeps mining | make up at 08:55; rpk group describe tayga-logminer at 09:00: 1 member, all 12 tayga.logs partitions, lag 0 to 4; log_templates FINAL: max(last_seen) 09:00:23 at 09:00:23, 69 templates seen in the last 2 min; again 09:11:36 at 09:11:37 |
verified live |
| Per-partition watermark keys and a per-replica heartbeat; the first assignment is seeded from the global key | logminer_state FINAL at 09:00: new_template_watermark_ns:p0 to :p11 (all 08:59:25.26), logminer_heartbeat_ns:b38d0e45749b (08:59:28); first assignment logged watermark 1791276937068679792, the value of the old global key new_template_watermark_ns |
verified live |
| Two replicas split the partitions 6 and 6 | LOGMINER_REPLICAS=2 make up at 09:11; rpk group describe at about 09:13: 2 members, partitions 0 to 5 on one, 6 to 11 on the other; logs: replica 01f7f4be63ab Assign([0..11]), Revoke, Assign([0, 1, 2, 3, 4, 5]), replica 3a32f1355256 Assign([6..11]) |
verified live |
| The two replicas mine disjoint services | shutdown log of each replica at 09:50:08: final new-template pass with services: 11 (partitions 6 to 11) and services: 7 (partitions 0 to 5, including tayga-e2e-probe); the partition map above has no service in both halves; alerts logged 09:32 to 09:44: replica 3a32… only frontend, payment; replica 01f7… only frontend-proxy, tayga-e2e-probe |
verified live |
| Both heartbeats are fresh with two replicas | logminer_state FINAL at 09:51:27: logminer_heartbeat_ns:01f7f4be63ab 09:51:12, :3a32f1355256 09:51:13 |
verified live |
| Log e2e scenarios pass with two replicas, one alert each, no duplicate | cargo test -p tayga-e2e --test scenarios -- --ignored --test-threads=1 --exact new_template_from_probe log_spike_on_payment_failure silence_alert_and_delivery (after make flags-reset): 3 passed in 526.2 s (spike alert after 291 s, new after 60 s, silence 170 s after enabling). log_alerts FINAL from 09:32: one spike for the payment template (e1e262d539fee41e), one new for the probe (cdc3405223acf55e), one silence (a769fdfe56f6c48a). No alert_id with more than one row after FINAL, no (template, kind) with two alert ids, and every alert id appears in one replica’s log only |
verified live (one run) |
Full make e2e with one replica after the scale back |
make e2e 10:14:39 to 10:30:36, then make flags-reset: 10 passed, 0 failed, 948.4 s. Story scenarios: adFailure 50 s, paymentFailure 70 s, paymentUnreachable 100 s, productCatalogFailure 35 s, intlShippingSlowdown 155 s (the shipping pre-check did not stop early); log scenarios: spike after 296 s, new after 60 s, silence 175 s after enabling. log_alerts FINAL from 10:14:39: one spike for payment, one new for the probe (dd7a44ca613f5a96), one silence (11be0fb369458544); no (template, kind) with two alert ids |
verified live (one run) |
docker compose … stop tayga-logminer stops every replica, and start starts them |
at 09:50:08 both containers Exited (0) within 1 s, each logged final new-template pass (when: shutdown) and tayga-logminer stopped; start brought both back |
verified live |
| Scale back to 1: the remaining replica takes all 12 partitions, from the minimum watermark | without a rebuild: LOGMINER_REPLICAS=1 docker compose … up -d --no-deps tayga-logminer at 10:13:59 removed replica 2 and kept replica 1 (Up 21 minutes, same member id); it logged final new-template pass (when: revoke, 11 services) and Assign([0..11]) with watermark 1791281607904267760 (10:13:27.90), the minimum of the stored keys (:p6–:p11 10:13:27.90, :p0–:p5 10:13:52.89). With LOGMINER_REPLICAS=1 make up at 09:51 (all recreated): watermark 1791280329344745000 = min(:p0–:p5 09:52:09.34, :p6–:p11 09:52:10.24) |
verified live |
| The single replica keeps mining every service after the takeover | log_templates FINAL at 10:14:32: max(last_seen) of every demo service between 10:14:09 and 10:14:32, including services of both former halves |
verified live |
| Docker DNS returns every replica | docker exec tayga-api getent hosts tayga-logminer six times: both addresses each time, in varying order |
verified live |
| The recorder stayed on one replica | metric_samples for tayga_logminer_logs_mined_total, 09:13 to 09:47: 138 samples, no decrease; the last (92,421) matches replica 3a32… (93,022 a few seconds later), not 01f7… (48,631); reqwest 0.13.5 pool_idle_timeout default 90 s (async_impl/client.rs) |
verified live; the 90 s default read in the crate source; superseded: plan 9 scrapes every replica (see the plan 9 rows) |
The Pipeline lag row reads the logminer on tayga.logs |
before the fix: tayga-logminer on tayga.signals lag 21,456 (09:02) and 469,682 (10:16), stale since the switch-over; after the fix (LOG_GROUPS in crates/tayga-api/src/lag.rs) and make up, GET /api/v1/pipeline/lag at 10:35 UTC: writer and assembler on tayga.signals 0, tayga-logminer on tayga.logs 0, tayga-notifier on tayga.alerts 0; live lag_it 2/2 |
verified live |
tayga.logs size |
rpk cluster logdirs describe --topics tayga.logs,tayga.signals --aggregate-into topic at 10:16: 70,614,087 bytes (first record about 08:56) and 25,382,696,990 bytes |
verified live; the 9 GB figure is an extrapolation |
| Stale heartbeat keys stay | logminer_state FINAL at 10:14: the global key (08:55:38) and keys of 4 replaced containers (b38d0e45749b, 01f7f4be63ab, 3a32f1355256, 0d2a9ee810c7) next to the live c0257fa435cf |
verified live; superseded: since plan 9 keys older than a day are deleted at startup |
| Container commands by compose service | docker compose … logs tayga-logminer printed the logminer’s logs (run with one replica; with two, stop and start were checked, see above); docker compose … exec --index 1 tayga-logminer bash -c '…/dev/tcp/127.0.0.1/9100…' returned tayga_logminer_logs_mined_total 80233; logs --index 1 --tail 1 worked; --index is in docker compose logs --help and exec --help (Compose v5.5.1). No docker exec/logs/restart tayga-logminer remains in the Makefile, README instructions, crates/tayga-e2e or deploy/ (grep); older Verified rows that name the container are history |
verified live |
LOGMINER_REPLICAS reaches compose |
export LOGMINER_REPLICAS ?= 1 in the Makefile: a test target printed 1 by default, 2 from the environment, 3 from make … LOGMINER_REPLICAS=3; docker compose … config tayga-logminer with LOGMINER_REPLICAS=2 shows deploy: replicas: 2 and no container_name |
verified |
| Ownership window 60 min; revoke pass bounded at 30 s, shutdown pass at 20 s; alert ids | ownership_window_min: 60 in LogminerSettings::default; REVOKE_PASS_TIMEOUT, SHUTDOWN_PASS_TIMEOUT in crates/tayga-logminer/src/main.rs; id_hex(&["new", template]), ["silence", template, since], ["spike", template, started_minute] in crates/tayga-drain/src/detect.rs |
verified in code |
| Rebalance flush, commit semantics and the crash window | on_rebalance, flush, shut_down docs in crates/tayga-logminer/src/main.rs; unit tests a_rebalance_during_pending_work_commits_only_flushed_offsets, a_clean_shutdown_flushes_then_runs_a_bounded_new_template_pass; ignored live tests a_template_first_seen_after_the_last_pass_is_announced_at_revoke / _at_shutdown in the same file |
verified in code and tests; a refused or late commit and a crash were not produced live |
| Performance, scale, or latency claims | none made | n/a |
Plan 9 (hardening)
Section titled “Plan 9 (hardening)”Rows below checked 2026-10-06 from 16:10 to 17:34 UTC on feat/plan-9-hardening at 9656d3a (plus this docs commit).
| Claim | How verified | Result |
|---|---|---|
make up deploys plan 9: ClickHouse recreated with the users.d mount, migration 12 applied |
make up 16:11:47 to 16:12:25 (log: opentelemetry-demo-clickhouse-1 Recreate); schema_migrations: version 12 applied 16:12:24; docker exec opentelemetry-demo-clickhouse-1 ls /etc/clickhouse-server/users.d lists grafana-readonly.xml |
verified live |
tayga-api has a healthcheck and becomes healthy |
docker inspect -f '{{.State.Health.Status}}' tayga-api: healthy at 16:12:50 (first poll, 25 s after the start), and again after each later make up; docker inspect -f '{{json .Config.Healthcheck}}': the bash /dev/tcp GET /healthz test, interval 10 s, timeout 3 s, start period 10 s, retries 3 |
verified live |
The services run as tayga, uid 10001 |
docker exec tayga-api id: uid=10001(tayga) gid=999(tayga) groups=999(tayga); docker exec tayga-writer id -u: 10001; Config.User is tayga for api, writer, ingest, assembler, notifier and the logminer |
verified live |
| Base images pinned by digest | docker/Dockerfile: node:24-bookworm-slim@sha256:d6aa754f…7b20, lukemathwalker/cargo-chef:latest-rust-1.98-trixie@sha256:94a624f8…3289, debian:trixie-slim@sha256:a29215f6…f11f, each with a pin-date comment; the image built from them in this make up |
verified in the file and by the build |
Unknown /api/ path answers JSON 404 |
curl -si localhost:8090/api/v1/nope: HTTP/1.1 404 Not Found, body {"error":"not found"} |
verified live |
ClickHouse reads carry max_execution_time 15 |
system.query_log after SYSTEM FLUSH LOGS: the API’s error_stories FINAL read at 16:13:29, user default, Settings['max_execution_time'] = 15 |
verified live |
query_timeout_secs default 15, 0 turns both bounds off; a stopped read (code 159) and a request 5 s past the limit answer 504 {"error":"storage timeout"}; non-API paths unbounded |
default_query_timeout_secs, the layer only when > 0 and saturating_add(SLACK_SECS) in crates/tayga-api/src/main.rs; SLACK_SECS = 5, api_timeout, is_clickhouse_timeout in timeout.rs; ApiError::into_response in routes.rs; unit tests query_timeout_defaults_to_15_s_and_0_turns_it_off, a_slow_api_request_is_a_json_504, fast_api_requests_and_other_paths_are_not_bounded |
verified in code and unit tests; no live timeout was produced (the code-159 match is unit-tested only) |
| Group examples stay inside the window | GET /api/v1/story-groups/2495695361138993645?since=1h at 16:13:29: 20 examples, all with ts_ns inside the last hour (True) |
verified live |
| Migration 12 marks every existing alert, so the upgrade republishes nothing | before the deploy (16:11:46): log_alerts FINAL 560, log_alert_publications absent; after (16:12:57): 560 marks, all published_at 16:12:24.2196, 0 alerts without a mark; SHOW CREATE TABLE: ReplacingMergeTree(published_at) ORDER BY alert_id, TTL 7 days |
verified live |
| No republish storm; new alerts get marks | the logminer’s own /metrics at 16:22:43 (10 min after the start): tayga_logminer_alerts_republished_total 0, tayga_logminer_commit_failures_total 0; alerts written after 16:12:25: 2, without a mark 0. After the make e2e run (new replica since 16:34:48) at 16:51:32: both counters 0; alerts written after 16:12:25: 13 (3 new, 6 spike, 1 silence from the e2e run), without a mark 0 |
verified live |
| Republish details (owned services, last 24 h, oldest first, at most 1,000, stop at the first failed send, after every pass) | REPUBLISH_WINDOW_NS, republish_unpublished after the detection pass, send_until_failure in crates/tayga-logminer/src/main.rs; Store::unpublished_alerts (ORDER BY version LIMIT 1000) in crates/tayga-store/src/logs.rs; tests republishing_stops_at_the_first_failed_send and the logminer IT a_stored_but_unpublished_alert_is_published_on_the_next_pass (logminer ITs 4/4 passed live on 2026-10-06) |
verified in code and tests; no unpublished alert occurred live |
| Stale heartbeat keys are deleted at startup; watermark keys are not touched | system.query_log: at each logminer start (16:12:25, 16:26:11 twice, 16:34:48) DELETE FROM logminer_state WHERE startsWith(key, 'logminer_heartbeat_ns') AND value < <start − 24 h> (1791216745008305840 for 16:12:25), no exception. The oldest key before the deploy was 7.27 h old, so nothing was old enough to delete: 8 heartbeat keys before, 12 after the three restarts (one per new container); new_template_watermark_ns* keys 13 before and after; the new replica’s key logminer_heartbeat_ns:861eb5c6be0d was written at 16:13:27.56 |
the DELETE ran live; the removal of an old key is covered by the store IT stale_heartbeats_are_deleted_and_other_keys_never (store_it 30/30 passed live on 2026-10-06), not observed live |
| Final flush bounded at 15 s, inside the 40 s grace with the 20 s pass | SHUTDOWN_FLUSH_TIMEOUT 15 s, SHUTDOWN_PASS_TIMEOUT 20 s and the sum test in crates/tayga-logminer/src/main.rs; stop_grace_period: 40s |
verified in code and tests |
Two replicas: each is scraped, labelled instance |
LOGMINER_REPLICAS=2 make up 16:25:43 to 16:26:13; at 16:28:18 metric_samples for job tayga-logminer, metric up, last minute: 172.26.0.19:9100 4 samples, 172.26.0.38:9100 4 samples (hostname -i in replica 1 and 2 printed these addresses); rpk group describe tayga-logminer: 2 members |
verified live |
| The Pipeline rate sums the replicas | each replica’s own tayga_logminer_logs_mined_total: 3528 and 3269 at 16:29:00.22, 4734 and 4496 at 16:29:59.85, together +2,433 in 59.6 s = 40.8/s; GET /api/v1/pipeline/series?metric=tayga_logminer_logs_mined_total&kind=rate&job=tayga-logminer&since=15m bucket 16:29: 41.63/s (2.0% apart) |
verified live (one bucket) |
| The data-lag gauge is per replica and the overview shows the slowest | replicas’ tayga_logminer_data_lag_seconds at 16:28:18: 1.61 and 1.80; at 16:31:12 the newest samples were 0.8097 (.19) and 0.4131 (.38) and GET /api/v1/overview?since=15m returned data_lag_secs 0.809661583, the larger |
verified live |
| Right after a scale-up the stale unlabelled lag counts for up to 300 s | at 16:30:36 the overview returned 2.104536977, the unlabelled sample of 16:25:54.99 (before the scale-up), while both replicas were below 1 s; at 16:31:12, past 300 s, it returned the replicas’ 0.81 | verified live |
Gauge merge rule: max for up, *_data_lag_seconds, *_templates, sum for the rest; DNS bounded at 2 s, IPv4 preferred, sticky labels |
max_over_instances in crates/tayga-api/src/series.rs; resolve_bounded, distinct, endpoints in recorder.rs; unit tests only_lag_template_and_up_gauges_take_the_largest_replica, a_hanging_lookup_falls_back_to_the_configured_url, ipv4_addresses_win_over_ipv6, once_several_replicas_are_seen_the_target_stays_labelled |
verified in code and tests |
Back to one replica after a tayga-api restart: samples are unlabelled |
make up at 16:34:28 to 16:34:50 (it recreated tayga-api); at 16:35:57 the last minute’s up samples for tayga-logminer: 4, all with an empty instance |
verified live |
Grafana’s ClickHouse datasource works as grafana |
LOGMINER_REPLICAS=2 make up-extras 16:31:25; /api/datasources/uid/tayga-clickhouse/health at 16:32:30: {"message":"Data source is working","status":"OK"}; POST /api/ds/query with SELECT count() FROM tayga.error_stories: one frame, value 30524, no error |
verified live |
grafana is read-only except max_execution_time |
as grafana on :18123: SELECT count() FROM tayga.error_stories returned 30524; CREATE TEMPORARY TABLE t (x UInt8): Code: 164. DB::Exception: grafana: Cannot execute query in readonly mode. (READONLY); max_threads=1: Code: 164 ... Cannot modify 'max_threads' setting in readonly mode. (READONLY); max_execution_time=5 with SELECT 1: 1 |
verified live |
Every dashboard works with the read-only user (no readonly = 2 needed) |
every target of the 4 provisioned dashboards (/api/search, /api/dashboards/uid/…) posted to Grafana’s /api/ds/query without credentials (anonymous Viewer), last hour, at 16:32:47: 39 targets (10 ClickHouse, 29 Prometheus), all status 200, 0 errors, no Code: 164; system.query_log for user grafana in those 3 minutes: 17 queries, the only exception the deliberate CREATE TEMPORARY TABLE |
verified live |
| Pipeline dashboard units, 11 error queries; Prometheus scrapes each replica | /api/dashboards/uid/tayga-pipeline: panels 1, 2, 3, 8, 9 units suffix: rec/s, suffix: rows/s, suffix: /s, suffix: items/s, suffix: logs/s; panel 6 has 11 targets; Prometheus /api/v1/targets for job tayga-logminer: 172.26.0.19:9100 up, 172.26.0.38:9100 up |
verified live |
| The “Consumer lag per group” panel matches Redpanda per topic | its three expressions on Prometheus /api/v1/query at 16:34:19: assembler 1379, writer 24, logminer 8, notifier 0; rpk group describe summed per topic at 16:34:19: assembler tayga.signals 1259, writer tayga.signals 173, logminer tayga.logs 8, notifier tayga.alerts 0 (writer and assembler lag move by hundreds within seconds). rpk’s TOTAL-LAG for tayga-logminer (2,428,210) also counts its stale tayga.signals offsets (2,430,077 at 16:34:19) |
verified live (one reading; same order of magnitude) |
Extras stopped without down |
docker compose … --profile extras stop tayga-grafana tayga-prometheus at 16:34:28; docker ps -a: both Exited (0) |
verified live |
Host run with TAYGA__RECORD_SECS=0 adds no up = 0 rows |
16:13:56 to 16:14:56: the host tayga-api on 127.0.0.1:18090 answered /healthz ok; its log has metric recorder disabled (record_secs = 0) once; metric_samples with metric = 'up' AND value = 0 in the last 2 minutes: 0 |
verified live |
record_secs 1 to 4 logs a warning |
(1..5).contains(&settings.record_secs) and the warning text in crates/tayga-api/src/main.rs |
verified in code; not run |
| Existing topics are 24 h since the hand change; the disk recovered | the maintainers ran rpk topic alter-config with retention.ms=86400000 at 16:14:28. rpk topic describe -c at 16:15:25: retention.ms 86400000 DYNAMIC_TOPIC_CONFIG on tayga.signals, tayga.logs, tayga.stories, tayga.alerts; rpk cluster logdirs describe … --aggregate-into topic: tayga.signals 7,395,359,304 bytes, tayga.logs 338,891,199, tayga.stories 17,381,468, tayga.alerts 683,991. Docker disk (docker run --rm alpine df /): 92% at 16:10:48 (86,009,892 of 93,709,644 KiB used + available), 91% at 16:12:25, 70% at 16:15:25, 71% at 16:58. Redpanda’s volume (docker system df -v): 28.8 GB at 16:11, 9.14 GB at about 16:40 |
verified live; the change itself was made by hand on the live stack |
The rpk command in the README |
rpk topic alter-config --help in the stack’s Redpanda: rpk topic alter-config [TOPICS...] --set key=value, --no-confirm disables the prompt |
syntax verified from the help; not run for this page (the live change was the maintainers’) |
New topics get 24 h; -1 unlimited; 0 and below -1 refused; existing topics not altered |
default_retention_ms 86,400,000, validate, new_topic in crates/tayga-kafka/src/lib.rs; tests new_topic_sets_retention_and_max_message_bytes, validate_checks_retention; create_topics only creates |
verified in code and tests; no topic was created live in plan 9 |
| The 2026-10-06 disk incident | live 2026-10-06: opentelemetry-demo-redpanda-1 exited at 13:54:22 with code 133, its log ending in a vassert backtrace (docker inspect, docker logs); last span in tayga.spans 13:54:21; docker inspect StartedAt of the Redpanda container: 14:19:16. The 99% disk and the 28 GB volume were read by hand during the incident |
verified live on 2026-10-06; the 99% reading was not recorded with its command and is not re-checkable now |
The mock binds 0.0.0.0 because of Linux host-gateway |
live 2026-10-06 in Docker Desktop’s Linux VM (127.0.0.1:18198 refused from a bridge container to 172.17.0.1, 0.0.0.0 answered HTTP/1.0 200 OK); doc comment on MockWebhook::start |
verified live once on 2026-10-06, not re-run for this page |
| Counters count accepted and committed work; a taken metrics port fails startup | record_flush and the help “Rows stored and committed” in crates/tayga-writer/src/metrics.rs; ingest help texts “in requests Kafka accepted” / “in accepted requests”; spawn_server with ? in the writer, assembler, logminer and notifier main.rs, the logminer’s before load_miner; tests a_taken_metrics_port_is_an_error_and_a_free_one_binds, http_conversion_counters_count_only_accepted_requests |
verified in code and tests; metric names seen live in metric_samples |
| Flag lock; shipping places its own orders (300 s); ad 300 s summed over its groups | FlagLock::acquire and its message, SHIPPING_TIMEOUT, ORDER_EVERY, AD_FAILURE_TIMEOUT in crates/tayga-e2e/src/lib.rs; wait_for_group_sum(… "kind=error&service=ad" …) and the shipping test in tests/scenarios.rs |
verified in code |
make e2e, run 1 |
16:37:59 to 16:50:53 UTC, then make flags-reset: 9 of 10 passed in 772.7 s. adFailure 130 s, log spike 160 s, new-template probe 60 s (warm after 0 s), paymentFailure 156 s, paymentUnreachable 100 s, productCatalogFailure 25 s, raw span counts and service map ok, silence 135 s after enabling. shipping_slowdown_produces_slow_story_blaming_shipping stopped at its pre-check: user_checkout_single n=49 p99=0.211 ok=0 (it needs 50; the demo checkout had recovered only at 16:03, and the payment scenarios just before it made 16:40 to 16:50 almost all errors) |
verified live; the pre-check refused correctly |
make e2e, run 2 |
17:22:30 to 17:32:59 UTC, then make flags-reset: 10 of 10 passed in 622.3 s. adFailure 55 s, log spike 145 s, new-template probe 60 s (warm after 0 s), paymentFailure 45 s, paymentUnreachable 105 s, productCatalogFailure 25 s, raw span counts and service map ok, intlShippingSlowdown 45 s (2 own orders placed; the second, 8d39dec87c831246ae64f63a68f48bd4, became a 5.1 s slow story blaming shipping, group 7169277777437214946, listed under kind=slow&service=shipping), silence 135 s after enabling. Afterwards the logminer’s alerts_republished_total and commit_failures_total were 0, and the 20 alerts written since the deploy all had marks |
verified live (one run each; single-run times, not latency guarantees) |
npm --prefix ui run e2e after the final-review fixes (story-flow walks the first spans, accepting the empty attributes state; Pipeline page has 10 charts with Commit failures) |
against the live app on :8090 rebuilt by make up with those fixes: first run 135 passed, 4 failed (pipeline.spec.ts still expected 9 charts); after correcting the count to 10, 139 passed, 28 skipped, 0 failed in 2.0 min |
verified live |
Playwright (npm --prefix ui run e2e) |
run 1, 16:13:37 to 16:15:17 (after the deploy): 137 passed, 28 skipped, 2 failed (pages.spec.ts pipeline and palette.spec.ts in the light project), both on a 503 from GET /api/v1/pipeline/lag during the hand retention change at 16:14:28; run 2, 16:52:00 to 16:53:57: 135 passed, 28 skipped, 4 failed, all story-flow.spec.ts (one per project): the home page’s first story group was then a demo accounting order-consumed slow story (432123300619415570) whose root span has 0 attributes, and the test expects attributes on the first span. The pipeline and palette tests passed in run 2 |
not clean: no run passed in full; the failures depend on live data or on the concurrent broker change and neither was traced to a plan 9 change |
| UI unit tests are stable | npm --prefix ui test three times in a row: Test Files 38 passed (38), Tests 500 passed (500) each time |
verified |
Component tests wait 5 s; make ui-e2e waits up to 5 minutes for data |
configure({ asyncUtilTimeout: 5_000 }) in ui/src/test/setup.ts, testTimeout: 15_000 in ui/vite.config.ts; READY_TIMEOUT_MS = 5 * 60_000 and the four checks in ui/e2e/global-setup.ts; both Playwright runs printed [global-setup] http://localhost:8090 has data after 0 s |
verified in code and live |
| Rust gates | cargo fmt --check clean; cargo clippy --workspace --all-targets -- -D warnings clean; cargo test --workspace: 497 passed, 0 failed, 62 ignored (sum of the test result lines); npm --prefix ui run lint and typecheck clean |
verified |
| Performance, scale, or latency claims | none beyond the single runs above | n/a |
Sub-project 4 (performance) and plan 10 (trace search)
Section titled “Sub-project 4 (performance) and plan 10 (trace search)”Checked 2026-10-07 (UTC times) on branch feat/sp4-performance; the last row (plan 10) carries its own date.
| Claim | How verified | Result |
|---|---|---|
| Control window before the deploy | logminer /metrics at 04:19:09: logs_mined_total 1,462,992, templates_created_total 0, templates 423; at 04:29:47: 1,486,991, 0, 423 (+23,999 lines, +0 templates in 10 min 38 s). log_alerts FINAL of the last hour at 04:19:15: no rows |
verified live |
make up deploys the cache with the default scalar backend |
make up 04:29:55 to 04:30:52 (image tayga:dev created 04:30:48); tayga-api healthy at the first poll (04:31:13); the logminer logged tayga-logminer consuming at 04:30:50.695 with restored: 423, "fingerprinter":"scalar"; /metrics: tayga_logminer_fingerprinter{backend="scalar"} 1 |
verified live |
| Cache hit ratio at least 99 %, no collisions, no resets | /metrics at 05:01:14, 30 min after the deploy: logs_mined_total 69,569, fingerprint_cache_hits_total 69,359, misses_total 210 (99.70 % hits), collisions_total 0, cache_resets_total 0 for full and generalised |
verified live |
| Template creation no higher than the control | templates_created_total 0 at 05:01:14 (control: +0 in 10 min 38 s); templates stayed 423 |
verified live |
No new alert for a template that existed before the deploy |
SELECT count() FROM log_alerts FINAL WHERE kind = 'new' AND started_at > toDateTime(T0) AND template_id IN (… first_seen < toDateTime(T0)) with T0 = 04:30:52: 0. The only alert since T0 was a spike of otelcol-contrib “Exporting failed. Will retry the request after interval.” at 04:31:52, the collector retrying while ingest restarted |
verified live |
| Templates hit after the deploy are the same as before it | per-service groupUniqArray(template_id) of log_template_hits, 30 min before 04:19:15 and from T0 to 05:01: 17 services and 66 ids each. One id only after: otelcol-contrib 13677393793404690769 “Could not inspect updated container”, first_seen 2026-10-03 16:45:53 (an existing template, hit once at 04:32:00 by the deploy’s container recreation, as at earlier deploys); one only before: ad “Transport failed”, last seen 04:23:22 |
verified live; no template was created after T0 |
| Hits track logs mined | Pipeline recorder, every minute 04:31 to 05:01: logs mined per minute = cache hits + misses per minute (2,053 to 2,527). log_template_hits 04:31 to 05:01 by log time: 65,210 rows = uniqExact(log_id) = logs rows in the same window; the recorder’s mined count over those minutes was 68,709 (ratio 0.949). The control before the deploy had the same ratio: 22,797 rows against +23,999 mined (0.950) |
verified live |
| The Pipeline page shows the two new charts with data (12 charts) | /pipeline?since=1h opened at 05:02:46 in Playwright: 12 captions and 12 canvases; “Logminer fingerprint cache … Latest: hits 39/s, misses 0.1/s.”; “Logminer batch time … Latest: p50 0.010 ms, p99 0.061 ms.” GET /api/v1/pipeline/series?metric=tayga_logminer_mine_batch_seconds&kind=q50 over 30 min: 30 points, 8.6 to 9.9 µs |
verified live |
| Real-data differential passes | the last hour of logs exported at 04:31:37 (the corpus command in docs/perf/sp4-performance.md, with LIMIT 200000): 130,378 lines, 17 services, 03:31:37 to 04:31:36; TAYGA_CORPUS=<file> cargo test -p tayga-drain --release --test differential --test fingerprint: 5 and 2 passed. With TAYGA_CORPUS=/nonexistent.jsonl the corpus test fails with open /nonexistent.jsonl, so the variable is read |
verified |
Kill switch LOGMINER_FINGERPRINTER=off |
LOGMINER_FINGERPRINTER=off make up 05:36:59 to 05:37:26; logged "fingerprinter":"off" at 05:37:24; /metrics at 05:39:33: fingerprinter{backend="off"} 1, hits 0, misses 4,911 = logs_mined_total 4,911; at 05:40:03: hits 0, mined 6,090 (+1,179 in 30 s). make up at 05:40:11 recreated the logminer only; logged "fingerprinter":"scalar" at 05:40:14; at 05:41:18 backend="scalar" 1, hits 2,331 of 2,431 |
verified live |
mine_batch time live, scalar against off |
mine_batch_seconds sum and count: scalar 0.10756 s over 13,470 records at 05:01:14 (8.0 µs per record, 5.16 logs per record); off 0.043595 s over 1,251 records at about 05:40:05 (34.8 µs) |
measured live over two different short windows; not a controlled benchmark |
| Service map CPU down at least 5× at a comparable refresh rate | system.query_log (OSCPUVirtualTimeMicroseconds) with a script sending GET /api/v1/service-map three at a time every 10 s (no UI tab was open): old code 04:19:42 to 04:29:42, 180 runs, 291.5 CPU ms per refresh (315 CPU s/h at 1,080 runs/h). Before the single-flight (hour to 05:31:03): 177 baseline runs, 71.3 ms per refresh, 4.1×. With the single-flight (2e09ac2, deployed 06:36:34; hour 06:41:25 to 07:41:25): 1,083 window runs (25.1 ms, 27.15 CPU s) and 60 baseline runs, one in each minute (270.1 ms, 16.21 CPU s): 43.4 CPU s/h, 40.0 ms per refresh, 7.3× against the live old code and 5.9× against the isolated estimate (235 ms per refresh). Staggered requests before the single-flight (05:38 to 05:52): 41.1 ms per refresh, 7.1× |
verified live, synchronised and staggered |
endpoint_stats and op_stats read at least 4× less per run |
system.query_log, 60 runs/h each: endpoint_stats 505.7 MiB per run in the hour to 04:19:15 (old code) against 18.3 MiB in the hour to 05:31:03 (new code), 27.7×; op_stats 1.32 GiB against 79.3 MiB, 17.0×. CPU: 70.9 → 15.2 and 98.5 → 31.2 CPU s/h |
verified live |
| The fingerprint cache paragraph: hits skip tokenising and the tree; a generalisation, a restore or 10,000 entries clear the service’s cache; collisions and non-ASCII or NUL bodies take Drain | add_fingerprinted, record, restore and CACHE_MAX_PER_SERVICE in crates/tayga-drain/src/drain.rs; fingerprint_body in fingerprint.rs; tests in tests/differential.rs and the drain.rs unit tests |
verified in code and tests |
logminer.fingerprinter values, default, gpu without the feature, unknown values; rebalance keeps the backend |
Backend in crates/tayga-logminer/src/backend.rs (and its tests); the default "scalar" and validate in main.rs; TAYGA__LOGMINER__FINGERPRINTER: ${LOGMINER_FINGERPRINTER:-scalar} in deploy/compose.tayga.yaml; the Dockerfile runs cargo build --release --workspace --bins (default features) |
verified in code; gpu startup not run live (not in the image) |
| The six metrics and the 12 Pipeline charts | crates/tayga-logminer/src/metrics.rs (histogram exponential_buckets(1e-6, 4.0, 10)); live /metrics above; ui/src/features/pipeline/model.ts has 12 charts |
verified in code and live |
parallel and gpu block the consume loop |
on_message (a plain fn called from the async consume loop) calls Miner::mine_batch synchronously; ParallelFingerprinter uses into_par_iter().collect_into_vec; the GPU poll waits up to POLL_TIMEOUT (5 s) |
verified in code |
| Batch-5 timings 2.44, 2.44 and 2.43 µs | docs/perf/sp4-performance.md, Fingerprint backends (criterion run, 2026-10-07) |
cited from the bench report, not re-run |
| Map health baseline: minute floor of the earlier of the window end and now, one cached entry, single-flight, a past window does not replace a newer one | health_baseline_end and HEALTH_BASELINE_SECS in crates/tayga-api/src/params.rs; HealthCache and health_baseline in repo.rs (tokio::sync::OnceCell per minute, no spawned task); unit tests concurrent_calls_for_one_end_load_once, a_failed_load_caches_nothing_and_keeps_the_error, an_aborted_leader_lets_a_waiter_load, an_older_end_never_replaces_the_slot_and_a_newer_one_does, a_window_ending_ahead_of_now_keys_the_current_minute and a_failed_health_baseline_caches_nothing; IT the_map_baseline_is_cached_per_minute_and_the_window_is_not; live: 60 baseline runs in 60 minutes (one per minute) under three synchronised requests every 10 s |
verified in code, tests and live |
Baseline queries use argMax instead of FINAL, with two edges |
baseline_with and its doc in crates/tayga-store/src/store.rs; trace_summaries engine ReplacingMergeTree(span_count) ORDER BY trace_id (system.tables, live); IT argmax_baselines_equal_final_over_duplicates (store ITs 32/32 passed on 2026-10-07) |
verified in code and live schema; ITs not re-run for this page |
| Bench and test commands in Developer commands | [[bench]] names in the four crates’ Cargo.toml, groups stages, fingerprint, cached in crates/tayga-drain/benches/mining.rs; TAYGA_CORPUS in crates/tayga-drain/tests/corpus/mod.rs; TAYGA_REQUIRE_GPU and no GPU adapter: skipped in tests/backends.rs; wgpu 30.0.1’s default features include metal, vulkan, dx12 and gles |
verified in code; benches not re-run for this page; GPU tests run on the Mac only |
| GPU tests on the Mac | TAYGA_REQUIRE_GPU=1 cargo test -p tayga-drain -p tayga-logminer --features gpu at about 04:25: 162 passed, 0 failed, 4 ignored |
verified |
make e2e with the cache |
06:04:26 to 06:18:22, then make flags-reset: 10 passed, 0 failed in 817.7 s. adFailure 85 s, log spike 250 s, new-template probe 60 s (warm after 0 s), paymentFailure 141 s, paymentUnreachable 90 s, productCatalogFailure 25 s, intlShippingSlowdown 40 s (own order f100cf643918fd481b2d0da684b5fc69, a 5.2 s slow story blaming shipping), silence 120 s after enabling, raw span counts and service map ok. adFailure, paymentFailure and paymentUnreachable printed “OVER the 60s target”, a soft target the tests do not assert |
verified live |
| Playwright with 12 Pipeline charts | npm --prefix ui run e2e 06:18:40 to 06:20:20 against the app rebuilt by make up: 139 passed, 28 skipped, 0 failed in 1.7 min; pipeline.spec.ts (expects 12 captions and 12 canvases) passed in all four projects; [global-setup] … has data after 0 s |
verified live |
| Rust and UI gates (sub-project 4) | before the deploy, about 04:21 to 04:25: cargo fmt --check and cargo clippy --workspace --all-targets -- -D warnings clean; cargo test --workspace 537 passed, 0 failed, 65 ignored (sum of the test result lines); cargo clippy -p tayga-drain -p tayga-logminer --features gpu --all-targets -- -D warnings clean; npm --prefix ui run lint and typecheck clean; npm --prefix ui test 39 files, 503 tests passed |
verified; re-run at 06:22 after the doc edits with the same counts |
Trace search reads the window with argMax instead of trace_summaries FINAL (plan 10), with a post-lookup that drops stale versions; the same result rows as FINAL except on span_count ties |
TRACE_SEARCH, TRACE_VERSION_SLACK_SECS (600 s), TRACE_SEARCH_EXTRA_ROWS (50), newest_versions and trace_span_counts with their docs (the edges) in crates/tayga-api/src/repo.rs; IT argmax_trace_search_equals_final_over_duplicates (repo ITs 11/11; merges stopped on its table), which fails with argMin, a slack of 0, no post-lookup, 0 extra rows, and an older-ts tie-break; 2026-10-07 10:16:53 UTC, old and new SQL back to back on the live tayga database (readonly=2, use_query_cache=0, use_query_condition_cache=0, 3 runs each, system.query_log), default 1h request: 251.0 MiB against 18.7 (main query) + 84.3 (post-lookup) = 103.0 MiB, 2.4×: the 4× bytes target is missed (the main query alone reads 13.4× less; the post-lookup reads most of the trace_id column), accepted by the maintainers on 2026-10-07 because correctness comes first; CPU median 477 against 146 ms (3.3×); 8 requests returned the same rows as FINAL, 3 of them (default, service, limit=500) in another order among equal ts; ec70b40f check (window 09:20 to 09:40, service=payment): FINAL 0 rows, main query 1 (the 2-span fragment), after the post-lookup 0, and the deployed API returned []; after make up both queries carry max_execution_time 15 |
verified in code, tests and live; numbers in docs/perf/sp4-performance.md, Plan 10 |
