A tester finishes your demo, shrugs, and says "yeah, it was fun." You thank them, they leave, and you have learned nothing. Meanwhile the build quietly knew that they died four times on the same ledge, skipped the crafting menu entirely, and spent eleven minutes in a room you designed for three — it just never wrote any of it down.
Telemetry is how you stop losing that information. Not an analytics dashboard with thirty charts nobody reads: a handful of events, written to a local file, that turn "it was fun" into "seven of nine testers never opened crafting." You can add it to a prototype in an afternoon, and it pays for itself in the first session.
Start from questions, not from events
The failure mode of instrumentation is logging everything. You end up with a 40 MB text file, no idea which column matters, and a vague guilt about not analysing it. Do the opposite: write down the three to five questions this build needs answered, then log the minimum that answers them.
For a combat prototype, the questions might be:
- Does anyone use the parry, or do they mash attack?
- Where do players die, and is it the same three metres every time?
- Does the second weapon ever get equipped?
Those questions imply about six events. Everything else is optional. Before the next test you throw half of them away and add new ones, because you will be asking different questions by then.
A useful habit: keep the question list beside the build in your playtesting module in GameDesignerX, so a session's raw numbers stay attached to the thing you were trying to learn. Six months later, a log file with no question attached is archaeology.
The ten events that earn their place
Across most projects, the same small set does the heavy lifting. Treat this as a starting menu rather than a spec:
| Event | Payload | Answers |
|---|---|---|
session_start |
build version, platform, settings | Which build produced this data |
level_enter / level_exit |
level id, elapsed time, exit reason | Pacing, where runs end |
player_death |
position, cause, level id, attempt number | Difficulty spikes |
checkpoint_reached |
checkpoint id, time since last | Progress funnel |
ability_used |
ability id, context | Whether mechanics are discovered |
item_equipped |
item id, replaced item | Whether choices are real choices |
menu_opened |
menu id, duration | Discoverability of systems |
objective_state |
objective id, state (given/complete/abandoned) | Quest clarity |
setting_changed |
setting, old, new | Accessibility and comfort needs |
session_end |
reason (quit/crash/finish), total time | Completion vs. abandonment |
Two rules make this data usable. First, every event carries the build version — mixing two builds in one spreadsheet has wasted more evenings than any bug. Second, every event carries a monotonic timestamp in seconds since session_start, so you can reconstruct a run as a sequence, not just a pile of counters.
Keep the schema boring
One line per event, one file per session, newline-delimited JSON:
{"t":184.2,"ev":"player_death","build":"0.4.2-a","level":"mine_02","cause":"fall","pos":[412,-88],"attempt":4}
{"t":190.6,"ev":"checkpoint_reached","build":"0.4.2-a","level":"mine_02","cp":"mine_02_mid","since_last":71.3}
No database, no service, no account. Write to the user's local app data folder, name the file session_<utc-timestamp>_<short-random>.jsonl, and have testers zip the folder and send it back. When you outgrow that, a single HTTP POST per session is the next step — not per event, which turns a flaky café Wi-Fi connection into missing data.
If you already run a build pipeline, stamp the same version string into your build tracking entry and into the log header. Matching them by hand is exactly the kind of chore that gets skipped at 1 a.m. before a demo deadline.
Read it as funnels and heatmaps, nothing fancier
You do not need statistics. You need two views.
The funnel counts how many sessions reached each checkpoint. Here is a real-shaped result from a hypothetical nine-tester session, the kind of table you can build in a spreadsheet in ten minutes:
| Checkpoint | Sessions reaching it | Median time to reach |
|---|---|---|
| Tutorial exit | 9 | 2m 10s |
| First combat | 9 | 4m 05s |
| Mine entrance | 8 | 9m 40s |
| Mid-mine checkpoint | 3 | 21m 15s |
| Boss room | 2 | 27m 50s |
The cliff between "mine entrance" and "mid-mine" is your whole to-do list. Five of eight players stopped in one stretch of level, and no amount of post-session conversation would have given you that number with confidence.
The heatmap plots player_death positions on a top-down screenshot of the level. Ten deaths scattered across a room is difficulty. Ten deaths inside a two-metre radius is a design bug — usually an unreadable gap, an off-screen hazard, or a ledge that looks grabbable and isn't. You can draw this by pasting coordinates into any plotting tool, or by drawing dots in your image editor over the level map you already have in your level documentation.
The two-minute silence check
One derived metric is worth computing on every session: the longest stretch with no gameplay event at all. Sort those gaps descending and look at the top five. Long silences are either a player reading something, a player lost, or an event you forgot to log. All three are worth knowing, and it takes one pass over the file to find them.
What telemetry will not tell you
Numbers say where; people say why. The funnel above tells you five testers stopped in the mine. It cannot tell you whether they were frustrated, bored, or confused about which door was the exit — and those three problems have completely different fixes. Pair every instrumented session with observation notes and one exit question ("what were you trying to do when you stopped?"), and log the pairing so the qualitative note sits next to the quantitative run.
Telemetry also lies in small samples in one specific way: it makes a single outlier look like a trend. One tester who spent 40 minutes in a shop because they were reading item descriptions out loud to you will drag a median of nine sessions noticeably. Look at the per-session list before you trust any average, and prefer medians and counts over means at this scale.
Privacy and consent, briefly
Even a local log file deserves care. Log game state, not people: no file paths, no usernames, no machine identifiers, no IP addresses, nothing you would be uncomfortable showing the tester. Tell testers what you record and how to find the files, and if you ship telemetry in a public demo, say so in the store page or an in-game notice and offer an off switch. Depending on where your players are, collecting identifiable data without a clear basis can put you in regulated territory — the simplest way out is to not collect it at all.
A checklist for your next build
- Three to five written questions this build must answer
- Six to ten events, each traceable to one of those questions
- Build version and session-relative timestamp on every event
- One
.jsonlfile per session in a folder testers can find - A ten-line script that turns a folder of logs into a funnel table
- Death coordinates plotted over the level map
- Observation notes attached to each session
- Findings that survive the numbers filed as issues, not as memories
The last line matters most. Instrumentation is only worth the afternoon if the output changes the build. When a funnel cliff turns into an issue with an owner and a milestone, telemetry has done its job; when it turns into a screenshot in a chat channel, you have built a very precise way of feeling bad.
Log less than you think you need, read it the same day, and delete the events that never answered anything.