Researchers have published what they say is the first benchmark in which general-purpose AI models drive a real car, and the headline result is that most of them could not get round a cone course in a car park.

The test

DrivingBench, posted to arXiv on 30 September 2026 by Aditya Ramabadran, Simon Mahns and Tobias Gessler, put four frontier vision-language models in command of a 2022 Toyota Corolla fitted with comma hardware running openpilot. The models were given three tools: one to see the car's camera frames and telemetry, one to command steering and speed, and one to stop.

The course was a short point-to-point route through a parking lot, marked out with small cones, with straight sections, bends, left and right turns and a marked finish area. Speeds were low throughout.

One design decision makes it harder than it sounds. The car keeps moving while the model thinks, and a new command replaces the one still running. Inference latency is part of the task rather than something the harness hides: the model must observe, act, notice it was wrong and recover, all while the car is travelling.

The results

Model Outcome
GPT-6 Astra Completed the course, on its second attempt
Claude Fable 5.1 Did not finish
GPT-5.6 Sol Did not finish
Grok 4.6 Did not finish

Each model got up to three attempts in a single conversation, in its vendor's own harness — Codex, Claude Code and Cursor. Astra was the only finisher, and the authors note that no other attempt got past 50% of the course. Of the 11 runs, eight ended in a crash, and eight failed to clear 11% of the route, coming apart at the first bend.

The individual failures instruct more than the tally. Grok read a gap between cones as a gate and drove through it after two commands. One GPT model invented a rule about what the cone colours meant and acted on it, having been explicitly warned otherwise. Several simply did not turn far enough. Astra's successful run cost $7.74 in inference, about four times the earlier half-course attempts.

The authors' summary is that frontier models are "frankly, not good at this".

Why this is a Tesla story

Because it is the control condition for an argument people have been having about FSD for two years.

Tesla's case is that driving needs a purpose-built end-to-end network trained on driving, running at camera frame rate in the car. The implied counter-case — that general multimodal models are getting good enough to make a driving-specific stack transitional — now has a number attached, and the number is eight crashes in eleven runs on cones in a car park at walking pace. That is not a verdict on where general models end up; it is a measurement of where they are, on a course with no other traffic, no pedestrians, no weather and no speed.

It is also a reminder of what the hardware was already doing. The car ran openpilot, the comma.ai system NHTSA opened an investigation into in September for the same stopped-vehicle crash pattern it investigated Autopilot for. openpilot drives that Corolla perfectly well on its own. The models were not competing with nothing — they were being handed the wheel of a car that already had a driving system in it, and mostly doing worse.

What it means in Europe

Nothing changes in a European car this week, but the benchmark lands in the middle of an argument European regulators are actively having.

The case Tesla and others make to EU authorities is that a trained, driving-specific network is auditable in a way a general-purpose model is not: you can characterise its failure modes, version it, and test it against a scenario catalogue. The TCMV's handling of the Dutch Article 39 request — a discussion in October, a vote no earlier than December — is a process built around exactly that kind of documentation.

DrivingBench shows why the distinction is drawn where it is. A model that invents a rule about cone colours after being told that rule does not exist is demonstrating the failure mode type approval handles worst: not a sensing error, but a confident wrong inference with nothing to catch it. European approval of a driver-assistance system is a bet that its mistakes are enumerable. On this evidence, general models' mistakes are not yet.