Why Rocketgraph?
One store, three signals
A trace, the logs it produced, and the metrics around it are joined by trace id — so you
go from a slow endpoint to the exact request and the log line that explains it.
Watchtower
Pick a minute, or the moment an alert fired, and get every trace, log and metric in that
window — grouped. Nothing infers a cause: the alert already said something was wrong,
and the log line already says why.
Pattern grouping
Drain3 template mining collapses hundreds of thousands of log lines into a handful of
patterns, so a new failure mode is visible instead of buried.
Runs in your cloud
Self-hosted by default. Data stays in your account and your ClickHouse — there is no
egress to a vendor.
How it works
- Send telemetry — Point any OpenTelemetry SDK or collector at your ingress endpoint with a bearer token. There is no proprietary SDK and no agent you have to trust.
- Everything lands together — Traces, logs and metrics are stored side by side and correlated by trace id, so a single query can cross signals.
- Investigate a window — Watchtower takes a time range and returns RED metrics per service, grouped log and trace patterns, the worst p95/p99, and the raw rows behind all of it.
- Alert on it — Define alert rules over the same data, and jump straight from a firing alert into the window it fired on.
What you get
- Distributed traces with a span waterfall, error roots and per-service timing
- Logs with server-side filtering, level and service facets, and full attributes per line
- Metrics from your services plus host metrics from the collector
- Infrastructure — AWS instances and CloudWatch metrics via a CloudFormation stack
- Alerting through Grafana rules over the same store
Next steps
Quickstart
Send your first span in under five minutes.
Node.js
Express, Next.js and NestJS.
Python
Django, Flask and FastAPI.
Java
Spring Boot and Quarkus, no code changes.
Kubernetes
Collector DaemonSet, or fan out from one you already run.