Back to Blog
telemetryplaytestingmetricsprototyping

Ten Events Worth Logging: Instrument Your Prototype Before the Next Playtest

"It was fun" tells you nothing. A handful of logged events turns a playtest into a funnel table and a death heatmap you can act on. What to log, how to store it, and what numbers will never tell you.

GameDesignerX TeamSeptember 28, 20267 min read

A tester finishes your demo, shrugs, and says "yeah, it was fun." You thank them, they leave, and you have learned nothing. Meanwhile the build quietly knew that they died four times on the same ledge, skipped the crafting menu entirely, and spent eleven minutes in a room you designed for three — it just never wrote any of it down.

Telemetry is how you stop losing that information. Not an analytics dashboard with thirty charts nobody reads: a handful of events, written to a local file, that turn "it was fun" into "seven of nine testers never opened crafting." You can add it to a prototype in an afternoon, and it pays for itself in the first session.

Start from questions, not from events

The failure mode of instrumentation is logging everything. You end up with a 40 MB text file, no idea which column matters, and a vague guilt about not analysing it. Do the opposite: write down the three to five questions this build needs answered, then log the minimum that answers them.

For a combat prototype, the questions might be:

  1. Does anyone use the parry, or do they mash attack?
  2. Where do players die, and is it the same three metres every time?
  3. Does the second weapon ever get equipped?

Those questions imply about six events. Everything else is optional. Before the next test you throw half of them away and add new ones, because you will be asking different questions by then.

A useful habit: keep the question list beside the build in your playtesting module in GameDesignerX, so a session's raw numbers stay attached to the thing you were trying to learn. Six months later, a log file with no question attached is archaeology.

The ten events that earn their place

Across most projects, the same small set does the heavy lifting. Treat this as a starting menu rather than a spec:

Event Payload Answers
session_start build version, platform, settings Which build produced this data
level_enter / level_exit level id, elapsed time, exit reason Pacing, where runs end
player_death position, cause, level id, attempt number Difficulty spikes
checkpoint_reached checkpoint id, time since last Progress funnel
ability_used ability id, context Whether mechanics are discovered
item_equipped item id, replaced item Whether choices are real choices
menu_opened menu id, duration Discoverability of systems
objective_state objective id, state (given/complete/abandoned) Quest clarity
setting_changed setting, old, new Accessibility and comfort needs
session_end reason (quit/crash/finish), total time Completion vs. abandonment

Two rules make this data usable. First, every event carries the build version — mixing two builds in one spreadsheet has wasted more evenings than any bug. Second, every event carries a monotonic timestamp in seconds since session_start, so you can reconstruct a run as a sequence, not just a pile of counters.

Keep the schema boring

One line per event, one file per session, newline-delimited JSON:

{"t":184.2,"ev":"player_death","build":"0.4.2-a","level":"mine_02","cause":"fall","pos":[412,-88],"attempt":4}
{"t":190.6,"ev":"checkpoint_reached","build":"0.4.2-a","level":"mine_02","cp":"mine_02_mid","since_last":71.3}

No database, no service, no account. Write to the user's local app data folder, name the file session_<utc-timestamp>_<short-random>.jsonl, and have testers zip the folder and send it back. When you outgrow that, a single HTTP POST per session is the next step — not per event, which turns a flaky café Wi-Fi connection into missing data.

If you already run a build pipeline, stamp the same version string into your build tracking entry and into the log header. Matching them by hand is exactly the kind of chore that gets skipped at 1 a.m. before a demo deadline.

Read it as funnels and heatmaps, nothing fancier

You do not need statistics. You need two views.

The funnel counts how many sessions reached each checkpoint. Here is a real-shaped result from a hypothetical nine-tester session, the kind of table you can build in a spreadsheet in ten minutes:

Checkpoint Sessions reaching it Median time to reach
Tutorial exit 9 2m 10s
First combat 9 4m 05s
Mine entrance 8 9m 40s
Mid-mine checkpoint 3 21m 15s
Boss room 2 27m 50s

The cliff between "mine entrance" and "mid-mine" is your whole to-do list. Five of eight players stopped in one stretch of level, and no amount of post-session conversation would have given you that number with confidence.

The heatmap plots player_death positions on a top-down screenshot of the level. Ten deaths scattered across a room is difficulty. Ten deaths inside a two-metre radius is a design bug — usually an unreadable gap, an off-screen hazard, or a ledge that looks grabbable and isn't. You can draw this by pasting coordinates into any plotting tool, or by drawing dots in your image editor over the level map you already have in your level documentation.

The two-minute silence check

One derived metric is worth computing on every session: the longest stretch with no gameplay event at all. Sort those gaps descending and look at the top five. Long silences are either a player reading something, a player lost, or an event you forgot to log. All three are worth knowing, and it takes one pass over the file to find them.

What telemetry will not tell you

Numbers say where; people say why. The funnel above tells you five testers stopped in the mine. It cannot tell you whether they were frustrated, bored, or confused about which door was the exit — and those three problems have completely different fixes. Pair every instrumented session with observation notes and one exit question ("what were you trying to do when you stopped?"), and log the pairing so the qualitative note sits next to the quantitative run.

Telemetry also lies in small samples in one specific way: it makes a single outlier look like a trend. One tester who spent 40 minutes in a shop because they were reading item descriptions out loud to you will drag a median of nine sessions noticeably. Look at the per-session list before you trust any average, and prefer medians and counts over means at this scale.

Privacy and consent, briefly

Even a local log file deserves care. Log game state, not people: no file paths, no usernames, no machine identifiers, no IP addresses, nothing you would be uncomfortable showing the tester. Tell testers what you record and how to find the files, and if you ship telemetry in a public demo, say so in the store page or an in-game notice and offer an off switch. Depending on where your players are, collecting identifiable data without a clear basis can put you in regulated territory — the simplest way out is to not collect it at all.

A checklist for your next build

  • Three to five written questions this build must answer
  • Six to ten events, each traceable to one of those questions
  • Build version and session-relative timestamp on every event
  • One .jsonl file per session in a folder testers can find
  • A ten-line script that turns a folder of logs into a funnel table
  • Death coordinates plotted over the level map
  • Observation notes attached to each session
  • Findings that survive the numbers filed as issues, not as memories

The last line matters most. Instrumentation is only worth the afternoon if the output changes the build. When a funnel cliff turns into an issue with an owner and a milestone, telemetry has done its job; when it turns into a screenshot in a chat channel, you have built a very precise way of feeling bad.

Log less than you think you need, read it the same day, and delete the events that never answered anything.