This is the read direction: Rocketgraph pulling from Datadog. If you want to
send telemetry to Rocketgraph instead, see
Instrumentation.
What it does
Every few minutes it pulls one window of logs, groups them into templates, and compares that window against its own recent history. When something has materially moved — and only then — it asks a model what it means.Step 1 — Create an API key
Go to Organization Settings → API Keys, or open app.datadoghq.com/organization-settings/api-keys. Click + New Key, give it a name you will recognise later (rocketgraph-read works), and copy the value. Datadog shows it once.
This key identifies your organisation. Rocketgraph uses it on read calls only —
it never submits data.
Step 2 — Create an application key, and scope it
This is the credential that actually grants access, and the one worth getting right. Go to Organization Settings → Application Keys, or open app.datadoghq.com/organization-settings/application-keys. Click + New Key, then set its scopes:
Nothing else.
Why two keys, and why only one can be scoped. The API key belongs to the
organisation and is what submits data, so it cannot be restricted. The
application key belongs to a user and is what authorises API reads, which is
why it supports scopes. If your security reviewer asks one question about this
integration, it will be this one.
Step 3 — Choose what to watch
Rocketgraph takes a Datadog log query — the same syntax as the Logs Explorer search bar.Step 4 — Connect it
In Rocketgraph, go to Datasources → Datadog and paste all three values.What it looks for
A real result
Six baseline windows on a steady system, then a fault appears:active=200 became active=<NUM>. Masking high-cardinality values is what
makes a template stable — without it, every line with a different number would
be its own pattern and nothing would ever group.
merchant=m_northwind was not masked, so it split into its own template.
That is correct: the literal recurs often enough to be a real pattern, and it is
exactly how a single bad merchant surfaces as a distinct signal instead of being
averaged into the rest.
And note the quiet windows say no model call. Nothing moved, so nothing was
asked.
First run is silent
On a cold start every template is technically new, which is true and useless. Rocketgraph reports nothing until it has history to compare against — typically six windows, so roughly half an hour at a five-minute cadence. After a deploy you will briefly see a burst of New templates. That is correct: new code emits new log lines. It settles once they are in the baseline.Troubleshooting
Every window says 'building baseline'
Every window says 'building baseline'
Snapshots are scoped by query text. If the query is changing between runs — a
shell variable, a trailing space, a different quoting — every run starts fresh.
Check they match exactly.
Volume appears to drop every single cycle
Volume appears to drop every single cycle
The window is probably ending at now. Log intake is not instantaneous, so the
most recent minute is always under-counted and reads as a fall. Rocketgraph ends
its window 60 seconds back for exactly this reason.
403 from Datadog
403 from Datadog
Almost always the site rather than the key. Check whether your organisation is
on
datadoghq.com, datadoghq.eu, or one of the US3/US5/AP1 deployments — the
site is shown in the top-right of Organization Settings.Nothing is ever reported, even during an incident
Nothing is ever reported, even during an incident
Check the query.
status:error on an estate that logs errors at warn will
match nothing. Run with * once to see what is actually there.What Rocketgraph does not do
- It does not write to Datadog. No monitors created, no tags added, nothing modified.
- It does not need an agent. No install on your hosts.
- It does not require migration. Keep sending to Datadog exactly as you do now.
- It does not call a model unless something changed. A healthy system costs nothing to watch.