Researchers have published what they say is the first benchmark in which general-purpose AI models drive a real car, and the headline result is that most of them could not get round a cone course in a car park.
The test
DrivingBench, posted to arXiv on 30 September 2026 by Aditya Ramabadran, Simon Mahns and Tobias Gessler, put four frontier vision-language models in command of a 2022 Toyota Corolla fitted with comma hardware running openpilot. The models were given three tools: one to see the car's camera frames and telemetry, one to command steering and speed, and one to stop.
The course was a short point-to-point route through a parking lot, marked out with small cones, with straight sections, bends, left and right turns and a marked finish area. Speeds were low throughout.
One design decision makes it harder than it sounds. The car keeps moving while the model thinks, and a new command replaces the one still running. Inference latency is part of the task rather than something the harness hides: the model must observe, act, notice it was wrong and recover, all while the car is travelling.
The results
| Model | Outcome |
|---|---|
| GPT-6 Astra | Completed the course, on its second attempt |
| Claude Fable 5.1 | Did not finish |
| GPT-5.6 Sol | Did not finish |
| Grok 4.6 | Did not finish |
Each model got up to three attempts in a single conversation, in its vendor's own harness — Codex, Claude Code and Cursor. Astra was the only finisher, and the authors note that no other attempt got past 50% of the course. Of the 11 runs, eight ended in a crash, and eight failed to clear 11% of the route, coming apart at the first bend.
The individual failures instruct more than the tally. Grok read a gap between cones as a gate and drove through it after two commands. One GPT model invented a rule about what the cone colours meant and acted on it, having been explicitly warned otherwise. Several simply did not turn far enough. Astra's successful run cost $7.74 in inference, about four times the earlier half-course attempts.
The authors' summary is that frontier models are "frankly, not good at this".
Why this is a Tesla story
Because it is the control condition for an argument people have been having about FSD for two years.
Tesla's case is that driving needs a purpose-built end-to-end network trained on driving, running at camera frame rate in the car. The implied counter-case — that general multimodal models are getting good enough to make a driving-specific stack transitional — now has a number attached, and the number is eight crashes in eleven runs on cones in a car park at walking pace. That is not a verdict on where general models end up; it is a measurement of where they are, on a course with no other traffic, no pedestrians, no weather and no speed.
It is also a reminder of what the hardware was already doing. The car ran openpilot, the comma.ai system NHTSA opened an investigation into in September for the same stopped-vehicle crash pattern it investigated Autopilot for. openpilot drives that Corolla perfectly well on its own. The models were not competing with nothing — they were being handed the wheel of a car that already had a driving system in it, and mostly doing worse.
What it means in Europe
Nothing changes in a European car this week, but the benchmark lands in the middle of an argument European regulators are actively having.
The case Tesla and others make to EU authorities is that a trained, driving-specific network is auditable in a way a general-purpose model is not: you can characterise its failure modes, version it, and test it against a scenario catalogue. The TCMV's handling of the Dutch Article 39 request — a discussion in October, a vote no earlier than December — is a process built around exactly that kind of documentation.
DrivingBench shows why the distinction is drawn where it is. A model that invents a rule about cone colours after being told that rule does not exist is demonstrating the failure mode type approval handles worst: not a sensing error, but a confident wrong inference with nothing to catch it. European approval of a driver-assistance system is a bet that its mistakes are enumerable. On this evidence, general models' mistakes are not yet.