Skip to main contentSkip to footer

Revisiting the Sensors in ADAS & Autonomous Vehicles (AV)

The discussion on the appropriate class of sensors (camera, radars, Lidars) is as old as the technology of driving assist and self-driving capability. While the majority of the industry leans on a multi-modal sensing, there is no shortage of proponents of the vision-only autonomous systems. The industry is split primarily into two schools or approaches. The bulk of the vehicle systems on the road today (including the current robotaxi leader Waymo) maximize redundancy using cameras with radars or a radar/Lidar combination, depending on the application (ADAS or AV).

There is a clear segment of proponents that support the notion that just like humans, who manage to navigate mainly with their vision, vehicles should also be able to safely navigate solely using their camera. The school of thought, primarily led by Tesla (see the discussion under Ref. 1), continues to make a case, despite several events attributed to camera’s limited sensing, an interesting one being the so-called ‘Coyote and the road runner’ trick in early 2025 (see Ref. 2), where a Tesla Model Y drove through a large Styrofoam wall with a realistic image of the road and the surrounding painted on it. Another event occurred earlier in 2022, when a Tesla Model 3, driven by the Smart Summon command to a private airport tarmac, failed to sense a small Cirrus Vision jet airplane parked in its way and ran into its tail, spinning the plane 180 degrees, while continuing to inch forward (see Ref. 3). While the jet was an unusual obstacle, and we are not privy to any analysis of the incident by Tesla or any reliable third party, there are speculations that the vision system failed to classify the jet as a hazard or could not perceive the tail as an obstacle requiring braking.

However, the camera detection algorithms have arguably made a lot of progress since the above incidents. It is also possible that even a radar might not have detected as thin a stationary structure as a small aircraft tail or the Styrofoam screen.

At the same time, it is also conceivable that a Lidar’s 3D point cloud could have helped in either scenario. This is one reason that the multi-modal sensor proponents have historically favored a combination of radar and Lidar, at least as a redundancy layer for detecting or confirming a variety of obstacles under various weather and light conditions.

The Sensor Tradeoffs

The industry is well aware of the relative pros and cons of tradeoffs, including the unit cost of cameras, radars, and Lidars, as shown below. The question today may be more about what combination is considered safe enough, what the cost-benefit ratio is, and the underlying risk-benefit equation. While Tesla remains the unapologetic flag-bearer of a camera-only system, most of the traditional OEMs and experienced robotaxi players, such as Waymo, favor a multi-modal sensor fusion. Others, such as Mobileye, promote both the camera-only system (see Ref. 4, Mobileye blog) and a hybrid variant that uses the camera as the primary sensor, enhanced by radar and Lidar as a confirmatory redundancy.

The table below summarizes the strengths and weaknesses of the sensors deployed in ADAS systems and autonomous vehicles.

The Case for a Camera-only System

Limited as it appears to be, vision-only sensing is not without its merits. As Andrej Karpathy, former head of Tesla AI, explained to MIT’s Lex Fridman in 2022 (Ref. 1), one of Tesla’s rationales for removing radar and ultrasonic sensors is that they add cost and a series of complexities, including manufacturing, supply chain, calibration, maintenance, and corresponding bloat in the software. Since Tesla’s approach has been to learn from real-world sensing data by using its entire fleet and build an ML/AI data engine, a multi-modal system (i.e., radar, Lidar, camera), as opposed to using a single modality (i.e., vision), would have been exponentially costly and complex. However, his argument and insistence that the camera is both “necessary and sufficient” is potentially problematic, given the obvious limitations in rain, hail, fog, and in low or zero-light situations.

A Key Challenge: High Dynamic Range (HDR) limitation

This refers to the physical limitations of conventional camera sensors to discern and capture the road situation under both extremely bright highlights and deep shadows simultaneously, or while transitioning from one extreme to the other, without losing critical visual details. To be truly effective, cameras should not only handle situations with poor light but also navigate environments like exiting a dark tunnel into blinding midday sunlight or identifying a dark-clothed pedestrian standing directly in the glare of oncoming headlights.

A Case for Enhancing Auto Safety with Thermal Cameras

Somehow, enhancing camera-vision’s HDR and darkness challenge with thermal capabilities has not really gained much traction, despite its proven potential. Thermal cameras have been in existence for decades, especially in the military, but were never seriously considered for automotive applications, except as a premium night vision feature in a few luxury models that never caught on. However, given the capability of thermal IR cameras to passively detect small heat signatures, they offer the new possibility of complementing RGB cameras to expand sensing of the surrounding environment at night or in other low-light or high-contrast visibility situations. They use Long-Wave Infrared (LWIR) cameras to detect blackbody radiation emitted from objects at terrestrial ambient temperatures (-23C to 77C or -9.7F to 170F).

A few years back, in 2024, VSI Labs, an ADAS analytics company, used a research vehicle at Michigan’s American Center for Mobility (ACM) to demonstrate the effectiveness of combined RGB and thermal cameras for pedestrian automatic emergency braking (PAEB) testing, conforming to FMVSS 127 standards. A broad PAEB function, effective at all hours, particularly at night, can be critical and effective in avoiding pedestrian fatalities, either at night or under high dynamic range situations, such as exiting a dark tunnel into a sunny exterior. According to NHTSA data, well over 70% of pedestrian fatalities occurred at night or in poorly lit conditions, with numbers still rising (Ref. 5).

The test, which I too attended, used thermally active pedestrian test mannequins (PTM) and a test vehicle with sensor fusion of optical and thermal cameras from Teledyne FLIR, comparing the test vehicle against four model year 2024 off-the-shelf vehicles from Tesla (Model Y), Subaru (Ascent), and Toyota (bZ4x).  While the test covered a gamut of use cases as per FMVSS No. 127 requirements, the key conclusion was that, while the off-the-shelf vehicles performed similarly to the test vehicle for daytime tests, all of them struggled for the nighttime tests, with the test vehicle being the only car that passed all the nighttime tests. Conclusions were similar for the high contrast use case (exiting a tunnel into bright sunlight).

Vision Sensing – Research Developments and New Enhancements

The potential for thermal cameras notwithstanding, there are new developments to improve the camera sensors, adding light-adaptive capability, i.e., instead of struggling when lighting conditions change, the technology automatically adjusts in real time, to maintain accuracy from bright daylight to dusk or near darkness. Penn State University researchers developed a breakthrough (see Ref. 6, 7, 8), inspired by the human retina, that tries to adapt and adjust when shifting from brighter to darker environments.

While still in a research state, the sensor upgrade has the potential to improve the role of vision sensing to better detect road objects more consistently, especially when visibility gets poor, an Achilles heel for cameras. This can also help avoid the loss of accuracy in glare, darkness, fog, or sudden changes in brightness.

Takeaway

The real-life challenges of road conditions under diverse weather, traffic, and light conditions demand a multi-modal approach for the foreseeable future. While computer vision-based neural networks may eventually reach (and hopefully exceed) human-level perception someday, the redundancy of combined multi-modal sensing and perception is still the safe and prudent approach, at least for now. For automotive-grade safety, we need the complementary capabilities of cameras, radars, and Lidars via smart sensor fusion – the redundancy adding to the confidence that a self-driving car will need to make the best decisions on the road.

However, the scaling of multi-modal sensing across the industry will need to improve the commercial economics of the sensors, especially the Lidars, which need to come well below the $100 threshold for wider adoption across the industry.

References:

Ref-1: Why Tesla removed Radar and Ultrasound sensors? Lex Fridman & Andrej Karpathy, https://www.youtube.com/watch?v=_W1JBAfV4Io&t=61s&ab_channel=LexClips (Lex Clips video series)

Ref-2: https://gizmodo.com/teslas-self-driving-fails-the-wile-e-coyote-test-2000577071

Ref-3: https://driving.ca/auto-news/crashes/watch-summoned-tesla-crashes-into-multi-million-dollar-jet

Ref-4: https://www.mobileye.com/blog/three-reasons-camera-first-adas-enables-scalable-automated-driving/

Ref-5: https://injuryfacts.nsc.org/motor-vehicle/road-users/pedestrians/

Ref-6: https://autos.yahoo.com/ev-and-future-tech/articles/self-driving-breakthrough-could-night-130000927.html

Ref-7: https://www.psu.edu/news/research/story/artificial-eyes-could-bring-human-sight-self-driving-cars-robots

Ref-8: https://www.dongascience.com/en/news/78322

#Sensors #Vision #Cameras #Radar #Lidar #autonomous #self-driving #HDR #sensorfusion #redundancy #ADAS

Next Post
Common Ground is not Common