Case study 03 · personal project · ADAS decision logic & verification tooling

openpilot's engage logic, ported and checked against my own car

My 2018 Toyota Prius, a hybrid, runs openpilot on a comma 3X. openpilot decides when it is allowed to steer with a five-state engage machine, written as nested Python if/elif. I ported that machine to a table-driven C++ library, built a layered, dependency-isolated replay stack around it, and checked the port against what the car actually logged: 5.5 hours of my own driving. This page covers the shape of both systems, how each rule traces to evidence, and where that evidence stops.

Aftermarket ADAS layer over OEM ECUs Pub/sub processes · cereal · msgq Table-driven FSM · pure core Differential replay C++17 · Cap'n Proto · zstd · GoogleTest Emscripten · WebAssembly
Type personal project, sole author Vehicle 2018 Toyota Prius (HEV) · openpilot on comma 3X Data 347 segments · 73 GB pulled from the comma 3X internal storage over adb
1,989,710
100 Hz ticks replayed, 5.5 h of real driving
99.9963%
per-tick state match, 74 mismatches
89 / 89
transitions in the same order, worst lag 20 ms
3 of 5
states ever seen on the road; the other two are test-only

01 — Architectural style, level 1: the vehicle

An aftermarket decision layer sitting on top of the OEM car

The car is two systems stacked. Underneath is Toyota's own vehicle: its ECUs, its hybrid powertrain, its CAN network. On top is openpilot, running on the comma 3X as separate processes that talk only through cereal, openpilot's messaging layer (msgq pub/sub, Cap'n Proto structs). The decision layer never touches the car directly. It publishes a command; a car-interface process turns that into CAN frames; the frames go out through a panda, which carries its own safety mode.

The engage state machine lives in selfdrived. It runs once per 100 Hz tick, publishes selfdriveState every tick, and publishes onroadEvents — the inputs it just acted on — once a second or on change. Those two topics are what my replay tool reads back from the log.

logical view — openpilot processes over the OEM networkanimated — one control loop, 10 s
L3 · DECISION · OPENPILOT PROCESSES ON THE COMMA 3X controlsdpubs carControl selfdrivedengage FSM · 100 Hz cardopendbc CarInterface ← selfdriveState · 100 Hz ← carState · 100 Hz carControl · 100 Hz → CEREAL · MSGQ PUB/SUB on-device logrlog.zst · one per minute onroadEvents · 1 Hz or on change my offloadadb pull → replay stack L2 · CAR INTERFACE · SAFETY pandadsets panda safety mode pandasafetyModel · controlsAllowed sendcan ↓ · can ↑ car CAN · OEM network OEM ECUOEM ECU L1 · TOYOTA'S VEHICLE hybrid powertrain stays Toyota's command state · events
Schematic. Process names, topics and rates come from openpilot's own source (selfdrived.py, card.py, controlsd.py, services.py); each process subscribes to more topics than drawn. The two ECU boxes stand in for Toyota's network, which I do not describe here. selfdrived also checks the panda: if the panda's safety mode does not match the car's config, or the panda has not allowed controls while openpilot is engaged, it raises an event that disengages.
What the "HEV" in the tab does and doesn't mean

The hybrid powertrain is Toyota's. I did not write, change, or reverse-engineer any of its control. My work here is on the decision layer above it: when openpilot is allowed to engage, and how to prove a copy of that logic behaves like the original. For production powertrain and supervisory control work, see the FCEV truck.

02 — Architectural style, level 2: the replay tool

A layered library stack, with the decision logic kept free of I/O

The tool is three static libraries and four front-ends. Each library depends only downward. oplog reads logs and is the only layer that knows about Cap'n Proto and zstd. opfsm is the state machine: a pure function of state, timer and events, with no Cap'n Proto, no I/O and no openpilot headers. opdiff is the only layer that knows about both, and it grades one against the other. The front-ends are thin.

structural view — libraries, front-ends, dependenciesanimated — what the browser build pulls in
L3 · FRONT-ENDS oplog-dumpwhat is in an rlog opdiff CLIgrade logs · exit code WASM APIextern "C" → browser fsm-diagramSVG · Mermaid L2 · COMPARE opdiffreplay · align · report L1 · READ | DECIDE oplogRlogReader · forEach(Event) opfsm16-row table · pure function L0 · DEPENDENCIES 5 vendored .capnp schemas+ stock libcapnp · libzstd nothingno capnp · no I/O · no openpilot
Taken from the CMake targets: opdiff links oplog and opfsm; fsm-diagram links only opfsm; the test binary links opfsm and opdiff. The WebAssembly build compiles the same three library sources. The schema compiler is a host binary, so the native build generates the Cap'n Proto sources and the wasm build reuses them.

At run time the stack is a straight pipeline. The one subtle part is timing. The log holds a state snapshot every tick but an event snapshot only once a second or on change, so the replay has to decide what the machine saw in between. It holds the last published set.

runtime data path — rlog.zst to reportanimated — pipeline and tick lane
rlog.zstbytes decompresszstd → u64 words parsecapnp Event walk eventsmask · held FSM stepone per tick comparevs logged state reportmatch · align oplog opdiff opfsm opdiff TICK LANE · SCHEMATIC SPACING selfdriveState · ~100 Hz · one FSM step per message onroadEvents · 1 Hz or on change · set held until the next publish on change
Stage ownership is from the source. The measured rates on comma's public demo segment were 97.3 Hz for selfdriveState and 1.6 Hz for onroadEvents. I first predicted that gap would wreck the match rate. It didn't: a change is exactly what triggers a publish, so the held sets land on the moments that matter. What it does cost is about one tick of lag.
ADR-02 · Do not link openpilot; vendor its schemasaccepted · builds in seconds, native and wasm
Context
openpilot's replay_lib would provide LogReader, FrameReader and route download. Setting it up pulls uv, git-lfs, six submodules, acados and scons, about 6 GB, and is developed and tested on Ubuntu 24.04. This project only reads log files.
Decision
Vendor five .capnp files, pinned by commit in VERSIONS.md, and link stock libcapnp and libzstd. Write log iteration myself, about 60 lines.
Rejected
Link replay_lib: its SConscript also links openpilot's common, messaging, visionipc and ffmpeg libraries, none of which a state-machine replay needs, and all of which would have to cross-compile for the browser.
Consequences
No migrateOldEvents(), so logs from before 2023 may not parse. Schema paths moved upstream in June 2026 (cereal/ became openpilot/cereal/), so re-pinning is a manual step.
ADR-03 · Keep going when a real-device log is truncatedaccepted · 15 of 346 segments
Context
Fifteen of the 346 replayed segments end mid-message, where the device lost power while writing. The first bulk run died 30 segments in, because one file's kj::Exception aborted the whole loop. The single-segment demo never hit this.
Decision
RlogReader::forEach stops cleanly at a partial trailing message and reports truncated. opdiff catches errors per file, prints SKIP <path>: <reason>, and carries on.
Rejected
Throw and abort: capnp framing is positional, so nothing after the break is reachable, but every whole minute before it is fine. Discarding those minutes throws away good evidence.
Consequences
The run reports 346 segments (0 unreadable, 15 truncated). That the corpus was messy is part of the result, not hidden. Exit status is still non-zero on any mismatch, so it can gate CI.

03 — Behavioural view

Five states, one timer, sixteen ordered rows

Upstream, the machine is 98 lines of Python in state.py. The states are disabled, preEnabled, enabled, softDisabling and overriding. It steps once per 100 Hz tick. Entering softDisabling starts a 3 s countdown, stored as 300 ticks (SOFT_DISABLE_TIME / DT_CTRL) so it cannot drift from the real one. From any engaged state, immediateDisable and userDisable go straight to disabled and outrank everything else.

The figure runs a copy of the 16-row table in the page. It steps at 100 Hz, and the fired row lights up below the chart. Play the scripted sequence, or toggle events yourself.

engage state machine — edges drawn from the tableinteractive — 100 Hz, row that fired highlighted
openpilot engage state machine Five states. From disabled, enable goes to enabled, enable plus preEnable goes to preEnabled, enable plus override goes to overriding, and noEntry loops back to disabled. preEnabled returns to enabled when preEnable clears. enabled goes to softDisabling on softDisable and to overriding on override. softDisabling returns to enabled when softDisable clears, loops while the timer is above zero, and times out to disabled. overriding returns to enabled when released, loops while override holds, and goes to softDisabling on softDisable. Four dashed red edges, one from each engaged state, go to disabled on immediateDisable or userDisable. When played, the current state and the edge that fired light up, and a readout shows the tick, events, fired row and soft-disable timer. enable +preEnable +override noEntry cleared preEnable softDisable cleared override released override softDisable timer > 0 timeout disabledpreEnabledenabled softDisablingoverriding tick0 events(none) fired(no row) timer0 / 300 immediate / user disable
  1. 1anyEngaged → disabled · immediateDisable
  2. 2anyEngaged → disabled · userDisable
  3. 3enabled → softDisabling · softDisable
  4. 4enabled → overriding · override
  5. 5softDisabling → enabled · softDisable cleared
  6. 6softDisabling → softDisabling · counting down
  7. 7softDisabling → disabled · timeout
  8. 8preEnabled → enabled · preEnable cleared
  9. 9preEnabled → preEnabled · preEnable
  10. 10overriding → softDisabling · softDisable
  11. 11overriding → enabled · override released
  12. 12overriding → overriding · override
  13. 13disabled → disabled · noEntry
  14. 14disabled → preEnabled · enable + preEnable
  15. 15disabled → overriding · enable + override
  16. 16disabled → enabled · enable
The rows, guards and 300-tick timer are the ones in opfsm/src/state_machine.cpp. Row 1 wins over row 16. The copy running in this page is a JavaScript mirror of that table, written for this figure; the scripted sequence is illustrative, not a logged drive. The edge set matches fsm-diagram's output. Node positions were placed by hand, as in the generated chart.
ADR-01 · Port the elif chain as an ordered transition tableaccepted · 16 rows · chart generated
Context
When the replay and the car disagree, I need to know which rule fired, and a nested if/elif can't say without a debugger. My hand-drawn chart for the first preview had 8 edges. The table has 16. The missing ones included the two highest-priority rules.
Decision
Each row is {from, to, trigger, guard, alerts, resets_timer}. The first row whose from matches and whose guard holds wins, so row order is the elif order. Step::fired points at the row that fired. States use cereal's ordinals, so logged state compares with no mapping. fsm-diagram draws edges from the same rows, so the picture cannot drift.
Rejected
Transliterate the elif chain: easiest to diff against upstream, but it can't answer "which rule?". Keep drawing the chart by hand: it looked right and was missing half the edges.
Consequences
Reordering rows changes behaviour, and the well-formedness test checks shape, not order. Display labels are shortened in the renderer, so the table stays about logic.

04 — Traceability

From logged behaviour back to the row that fired, and on to its evidence

Every rule has a chain: the upstream Python branch, the C++ row that copies it, the unit tests that pin it, the replay against logged selfdriveState, and a recomputation in the browser. The chain can be walked either way. Below it runs backwards, from something the car did to the line of state.py that explains it.

trace — one engage transition, logged state to upstream ruleanimated — 16 s
1 · THE CAR LOGGED disabled → enabled 2 · THE REPLAY TOOK THE SAME EDGE, IN THE SAME ORDER 3 · Step::fired POINTS AT TABLE ROW 16 4 · ROW 16 IS THE FINAL else IN state.py 5 · SAME ROW: PINNED BY TESTS, RECOMPUTED IN THE BROWSER loggedpredicted disabledenableddisabledenabled Upstream ruleC++ table rowUnit testsDiff replayBrowser (wasm) state.pyrow 16 of 1616 · GoogleTestvs logged statesame C++ if events.contains(ENABLE): ...no NO_ENTRY, no PRE_ENABLE, no override else: state = enabled {kDisabled, State::enabled, "enable", gEnableAllowed, ET_ENABLE} EngagementRoutes FromDisabledFiredTransitionIs ReportedForTheChart 5.5 h · 1,989,710 ticks89/89 transitionsin orderworst lag 20 ms 6,000 ticks6/6 alignedself-check:identical Reading right to left is debugging. Reading left to right is verification. The row sits in the middle of both.
One edge, traced for illustration. The band spacing is schematic. The replay reports whole transition sequences and lag, not how often each row fired, so the corpus numbers are evidence for the aligned sequence as a whole. The browser figures are from one segment of my driving, recomputed from its raw log in a local preview, and checked against the native run.

Behaviour ↔ rule ↔ evidence

BehaviourUpstream branchRowsUnit testsReal-drive replay
Immediate / user disable beat everythingouter if IMMEDIATE_DISABLE / elif USER_DISABLE1, 2ImmediateDisableFromEveryState, UserDisableFromEveryStateexercised: with soft disable never seen, only these rows lead from engaged to disabled
Engagedisabled: ENABLE branch15, 16EngagementRoutesFromDisabled, FiredTransitionIsReportedForTheChartexercised
Override enter, hold, releaseenabled elif OVERRIDE_*; overriding branch4, 11, 12OverrideReleasedReturnsToEnabled, MaintainStates. No test isolates row 4.exercised: 2,318 overriding ticks
noEntry refuses engagement and raises its alertdisabled: if NO_ENTRY13NoEntryBlocksEngagement, NoEntryRaisesItsAlertinvisible: replay compares state only, and this row does not change state
preEnable pathdisabled PRE_ENABLE; preEnabled branch8, 9, 14EngagementRoutesFromDisabled, NoEntryDoesNotKickOutOfPreEnabled, MaintainStatesnever occurs; unit tests only
Soft disable and 300-tick timeoutenabled / overriding SOFT_DISABLE; softDisabling branch3, 5, 6, 7, 10SoftDisable, SoftDisableTimerRunsOutAfterThreeSeconds, ClearingSoftDisableReturnsToEnabledEvenWithTimerExpired, SoftDisableFromOverridingRestartsTimernever occurs; unit tests only
The table as a wholewhole update()1–16TransitionTableIsWellFormed; chart generated from the rowsstates seen: disabled 1,771,242 · enabled 216,150 · overriding 2,318 ticks
Where the evidence stops

preEnabled and softDisabling never occur in 5.5 hours of my driving. The soft-disable countdown is the highest-consequence branch in the machine, and it has been checked only by unit tests, never against a real car. More miles of the same driving won't close that gap. It needs a drive where openpilot disengages itself, and that isn't something to stage on a public road. Two approximations also belong in the claim: event sets are held between publications, and a replay seeded mid-log assumes a full soft-disable timer, because the countdown isn't published.

The test that could not fail

All 16 tests passed after I deleted the noEntry row. Weakening its guard didn't fail them either. The enable guards also refuse engagement, and every test checked state, not alerts. Adding NoEntryRaisesItsAlert made the deletion fail the suite, 1 of 16.

An upstream test bug

In test_soft_disable, assert a == b if cond else State.softDisabling parses as (a == b) if cond else …. For every non-disabled state it asserts a truthy enum and cannot fail. The C++ port writes the comparison as intended, and that behaviour does hold.

Two replay modes

Free-running seeds from the log once and lets errors cascade, which is the honest measure. --resync snaps back after each mismatch. They report the same result on the demo route, which shows that nothing cascaded.

ADR-04 · Recompute in the browser from the same sourcesaccepted · 230 KiB module · 133 ms
Context
A page that states "99.9963%" is asking to be believed. A page that recomputes the result from a raw log with the same code gives the reader something to check.
Decision
Compile oplog, opfsm and opdiff sources with Emscripten, unchanged. Expose flat extern "C" functions, and pass results back as typed arrays over the heap. Build single-threaded, with exceptions on. The page compares its result with the native run and prints identical or DIFFERS.
Rejected
embind: tens of KB of glue for an interface that is a few scalars and two float arrays. pthreads: needs SharedArrayBuffer, which needs COOP/COEP headers that GitHub Pages cannot set. At 133 ms there is nothing to parallelise. A JS reimplementation: it would prove only the JS.
Consequences
Sharing sources makes Linux-only assumptions break loudly. That is how the std::vector<capnp::word> bug surfaced: libc++ rejected it and libstdc++ had allowed it. It needs http, not file://. Shipping a whole rlog works for one segment but not for a route.

05 — Contrast

Four ways to represent the same engage logic

The rules are fixed; they're openpilot's. The architectural choice was how to represent them so they can be run, drawn, diffed against upstream and explained when something disagrees. I built B. C and D are the options I'd weigh on a different program.

A · as upstream

Nested if / elif

Line-for-line with state.py. Priority is implicit in nesting, and "which rule fired" needs a debugger.

B · chosen

Ordered transition table

Priority is row order. Step::fired names the row. The chart is generated from it. Plain C++17 that builds for wasm.

C · model-based

Stateflow chart → generated C

The chart is the source; code is generated. Execution order is explicit on the chart, and model coverage tooling is available. It needs a MATLAB toolchain.

D · formal

Statechart spec, model-checked

A spec such as SCXML or a model-checker input, plus properties checked over every input sequence, not only the ones driven.

CriterionA · if / elifB · table (chosen)C · StateflowD · formal spec
Priority is explicit and reviewableimplicit in nestingrow orderexecution order on chartin the spec
"Which rule fired?" at replay timedebuggerStep::firedneeds instrumentation or loggingdepends on the runtime; a checked model does not run in the car
Diff against upstream Pythonline for linebranch to rowgraphical; hard to diffre-expression; farther from source
Diagram cannot drift from logicnogenerated from rowsdiagram is the sourcecan be generated
Covers soft disable without road dataunit testsunit tests (the gap above)tests plus model coverageproperties over all inputs
Toolchain cost · runs in a browser tabPython; not this stackno dependencies; wasmlicensed toolchain; generated C would still need porting inextra tool; spec kept beside the code
When I would choose differently

C, if this logic were being engineered rather than copied: an OEM or supplier program that already works in Simulink, reviews charts and generates production code. That is the toolchain I used on the FCEV program, and the table's shape (ordered, guarded, priority visible) maps onto a chart directly. D, added beside the implementation rather than instead of it, if the soft-disable gap mattered for release. It's the one option that can check "immediateDisable always reaches disabled in one step" or "softDisabling always ends within 300 ticks" without a road drive that no one should stage.

A smaller contrast: how the tool reads logs

link upstream

openpilot replay_lib

LogReader, FrameReader and route download included. About 6 GB of setup; also links common, messaging, visionipc and ffmpeg.

chosen

Vendor 5 schemas + stock libs

Log iteration rewritten (about 60 lines). No schema migration for old logs. Builds anywhere, including wasm.

CriterionLink replay_libVendor schemas (chosen)
Setup footprint~6 GB; uv, git-lfs, six submodules, acados, sconsfive text files, two stock libraries
Old logs (migrateOldEvents)yesno; pre-2023 logs may not parse
Video frames, remote routesincludednot provided
Same sources compile to wasmnot attemptedyes

I'd link replay_lib if the tool needed to decode video frames, fetch routes from comma's servers, or read logs from before 2023. For state and events alone, that dependency cost more than it returned.

Scope, stated plainly: this is a personal project on my own car, not job work. B and the vendored-schema reader are what I built and measured. A is upstream's code, which I read and ported. C and D are an analysis of alternatives; I did not build this logic in Stateflow or model-check it. openpilot, the panda and the car interface are comma's work; the powertrain and its ECUs are Toyota's. The upstream test bug above is found, not yet fixed upstream.