Regulators on both sides of the Atlantic are converging on a new premise: the safety of autonomous systems cannot be inferred from how broadly they are deployed. What matters is how they behave at the edges—boundary competence, in the technical vernacular. The problem is that no one has finished writing the standards for measuring it.

This week, the National Highway Traffic Safety Administration escalated its investigation of Tesla's Full Self-Driving (FSD) system from a preliminary evaluation to an engineering analysis covering approximately 2.9 million vehicles. The shift is significant. NHTSA is no longer content to study aggregate incident rates; it is examining whether camera-only perception can reliably detect hazards in fog, glare, and low-light conditions. Documented failures—including FSD v14 Lite driving over downed trees obscured by fog and running red lights—point to reproducible gaps in boundary recognition that cumulative mileage cannot surface.

The logic of the probe is worth parsing. Traditional automotive safety assumes that if you drive enough miles without incident, the system is probably sound. NHTSA is signaling that this assumption collapses when the failure mode is rare but catastrophic, and when the system's architecture lacks the sensory redundancy to perceive the hazard in the first place. A camera that cannot see through fog does not get safer with more miles; it merely accumulates more evidence of the conditions it handled correctly.

Meanwhile, Waymo is expanding operations into new U.S. cities and planning a 2026 launch in Germany, buoyed by third-party data showing lower crash rates than human drivers. Yet the company is simultaneously under federal investigation for incidents including passing a stopped school bus and a collision with a child. The pattern suggests that even bounded operational design domains and published safety data do not prevent incident-driven enforcement. When something goes wrong visibly, regulators respond regardless of aggregate statistics.

In Europe, the EU AI Act's high-risk obligations took effect on August 2, requiring risk management systems and transparency for automated systems operating in critical infrastructure. But CEN-CENELEC Joint Technical Committee 21, tasked with developing harmonized technical standards, has not finalized any. Companies must currently self-assess against legal language like "adequate" risk management—a phrase that pushes compliance toward legal interpretation rather than engineering verification.

The convergence is striking. Both jurisdictions are moving toward a regulatory posture where deployment breadth no longer substitutes for depth of validation. But without published sensor-performance mandates or harmonized EU standards, operators face a compliance gap: they are expected to demonstrate boundary competence without agreed-upon metrics for what competence entails. The risk is that enforcement becomes arbitrary and reactive, shifting incentives from structured verification to narrative management unless standards can catch up to investigations.

I do not know whether NHTSA's engineering analysis will yield a camera-performance threshold or whether the EU will publish harmonized standards before the end of the year. I do know that the question—how do we verify what a system cannot yet see—is now central to both. And I suspect that whoever answers it first will shape the template for autonomous system regulation for the next decade.

The miles will continue to accumulate. Whether they still matter depends on what we learn to measure.


Sources: