Demo with the OpenTelemetry demo
The OpenTelemetry demo is a microservice shop (checkout, payment, shipping, ad, product catalog and more) with a load generator that keeps buying things. Its services carry feature flags that inject failures. The Tayga repository vendors it as a git submodule, pinned to 3.1.0, and runs Tayga next to it: the demo’s Collector forwards every trace and log to Tayga as well as to the demo’s own backends.
This is the quickest way to see what Tayga is for. Plan on about 10 minutes from clone to your first injected failure as a story, plus the time of the first image build.
What you need
Section titled “What you need”- Docker with the Compose v2 plugin, git and make.
- Memory. The demo is large. On 2026-10-07 the running stack (the demo plus Tayga, 36 containers) used 8.9 GiB of memory by
docker stats, ClickHouse 2.9 GiB of it, on Docker Desktop with 14 CPUs and 31.5 GiB. Give Docker comfortably more than that. - Disk. Tayga keeps raw spans and logs for 3 days and its Redpanda topics for 24 hours; on the demo,
tayga.signalsheld 7.4 GB at 24-hour retention. See Retention and disk. - Optional: a Rust toolchain (1.98 or newer) for
make flagandmake e2e, which run thetayga-devtoolsCLI with Cargo. Without Rust, flip flags in the demo’s own flag page instead (step 4).
Run it
Section titled “Run it”-
Clone with the submodule.
Terminal window git clone --recurse-submodules https://github.com/softberries/tayga.gitcd taygamake upalso runsgit submodule update --init, so a clone without--recurse-submodulesworks too. -
Start everything.
Terminal window make upIt initialises the submodule, builds the
tayga:devimage (the web app is built inside it, so no Node is needed on the host), and starts the demo and Tayga together. The first build takes several minutes.tayga-migratecreates the ClickHouse schema first; the other Tayga services wait for it.Check that the API is up; it reads
healthyonce it serves:Terminal window docker inspect -f '{{.State.Health.Status}}' tayga-api -
Open the two apps.
What Where Tayga http://localhost:8090The demo shop http://localhost:8080The demo’s flag page http://localhost:8080/featureThe demo’s Jaeger (Tayga’s “Open in Jaeger” links go here) http://localhost:8080/jaeger/uiThe load generator starts at once, so traces and the service map fill as soon as the first traces close. Slow stories need a baseline first: an endpoint is judged against its own p99 of the last hour.
-
Inject a failure. Make every payment fail:
Terminal window make flag NAME=paymentFailure VARIANT=100%It sets the flag’s default variant in
deploy/flagd/demo.flagd.json, the file the demo’sflagdreads. The first run compilestayga-devtools.Open
http://localhost:8080/feature, set the default variant of paymentFailure to 100%. The page edits the samedeploy/flagd/demo.flagd.json. -
Watch it become a story. On Stories (keep the time range at 1h and Live on), a new group appears whose summary starts
payment charge failed: Payment request failed. Invalid token.In our end-to-end runs the first matching stories arrived 55 to 125 seconds after the flip (single runs, not a guarantee). The endpoint is part of the fingerprint, so you will see one group per calling endpoint, such asload-generator user_checkout_singleanduser_checkout_multi.Stories on the demo stack with paymentFailure on: the two “payment charge failed” groups at the top, one per calling endpoint. The inspector on the right shows the selected group's sample story (here a product-catalog failure). -
Open the story. Select the group’s sample story. The root-cause card names
paymentand itschargespan, with the path throughcheckout; the waterfall marks the failing chain and the critical path; the logs below are the request’s own, with thePayment request failederror first.A paymentFailure story: root cause, request path, group trend, comparison with normal, waterfall and logs. -
Look at the service map. Press g then m.
checkoutandpaymentturn red with their error ratio, and the call between them is drawn dashed as failing. Clickpaymentfor its rate, errors and p99, and its recent stories.The service map while payments fail: checkout and payment are degraded, and the failing call is dashed. -
Wait for the log alert. The failing payments log the same error again and again, so its log template spikes. Detection runs every 60 seconds over a 5-minute window, so the alert takes a few minutes: 175 to 291 seconds in our runs. Open Logs → Alerts (g then l). Each alert links example traces to their stories. With a notifier target set, it would also go to a webhook or Slack.
A spike alert: the count in the window against the baseline, and example traces linked to their stories. Here the demo's proxy logged the failing checkouts. -
Put the flags back.
Terminal window make flags-resetIt copies the demo’s default flag file over
deploy/flagd/demo.flagd.json, which turns every injected failure off. The group stops growing, its trend line flattens, and the alert stops being active about 10 minutes after the last spike.
Other failures to try
Section titled “Other failures to try”Each flag below is one of the demo’s own. Tayga’s end-to-end tests flip the same flags and check the outcome against the API; the times are from single runs between 2026-10-04 and 2026-10-06, not guarantees.
| Flag and variant | What Tayga shows | First stories after |
|---|---|---|
paymentFailure 100% (also 10% to 90%) |
Error stories blaming payment’s charge span, one group per calling endpoint; later a spike alert on the payment error template |
55–125 s; alert 175–291 s |
paymentUnreachable on |
Error stories whose root cause is checkout’s client span: “checkout could not reach oteldemo.PaymentService” |
50–80 s |
productCatalogFailure on |
Error stories in product-catalog: “Product Catalog Fail Feature Flag Enabled” |
25 s |
adFailure on |
GetAds failed stories in ad, one group per calling endpoint. The flag hits about one request in ten, so it is slower to show |
75–85 s |
intlShippingSlowdown 5sec |
A slow story blaming shipping, with the delay on the critical path. Only international orders are delayed, and they are rare in the load generator’s traffic; the end-to-end test places its own. It also needs a clean baseline: an hour without earlier slowdown runs |
40–160 s when it passed |
make flag NAME=paymentUnreachable VARIANT=onmake flags-resetRun the end-to-end tests
Section titled “Run the end-to-end tests”The same scenarios run as tests against the live stack. They take the flags one at a time and reset them before they start; full runs took 8 to 16 minutes.
make e2eDetails of each scenario and its timeouts are in From source.
Grafana and Prometheus (optional)
Section titled “Grafana and Prometheus (optional)”Tayga needs neither: its own Pipeline page records the services’ metrics. To add Tayga’s Grafana dashboards and a Prometheus next to the demo:
make up-extrasGrafana is then at http://localhost:3001 (anonymous Viewer access; the admin password is admin) and Prometheus at http://localhost:19090, and the service map gains an “Open in Grafana” link. See Metrics and Grafana.
Tayga’s own published ports bind to 127.0.0.1:
| Port | Service |
|---|---|
| 8090 | tayga-api: web app, JSON API, /healthz, /metrics |
| 14318 | tayga-ingest OTLP/HTTP (container port 4318), used by tayga-devtools emit-log |
| 18123 | ClickHouse HTTP (container port 8123) |
| 19092 | Redpanda Kafka API (external listener) |
| 3001, 19090 | Grafana and Prometheus, only after make up-extras |
| 8080 | The demo’s frontend proxy: the shop, /feature, /jaeger/ui (published by the demo) |
What Tayga changes in the demo
Section titled “What Tayga changes in the demo”The demo itself is not edited; Tayga’s Compose file (deploy/compose.tayga.yaml) layers a few changes on top of it:
- The Collector loads
deploy/otelcol-config-tayga.ymllast, which adds theotlp_grpc/taygaexporter to the traces and logs pipelines next to the demo’s own exporters. - The flag file is
deploy/flagd/demo.flagd.json, mounted intoflagdand its UI, somake flagandmake flags-resetcan change it. - The load generator skips its
ask_agenttask. It calls anagentservice that only exists in the demo’scompose.agent.yaml, which the Makefile does not include, so every call failed and showed up as noise. A small locustfile,deploy/locust/tayga_locustfile.py, imports the demo’s own and removes that task. - Memory limits of demo services that ran at their caps during long sessions are raised:
checkoutandproduct-catalog20M to 64M,adandfraud-detection300M to 512M,accounting160M to 320M,kafka620M to 1G,opensearch1G to 1.5G, the demo’sgrafana175M to 256M, andload-generatorto 2G.
Day to day
Section titled “Day to day”make ps # status of every containermake logs SERVICE=tayga-api # follow one service's logsLOGMINER_REPLICAS=2 make up # rescale the logminermake down # stop and remove the containers; the data volumes staymake up again after git pull rebuilds the image and migrates the schema before the services restart. If the demo misbehaves (every checkout failing for hours, say), Troubleshooting has what we have seen and how it was fixed.
