Thursday, 30 July 2026 Login

Spreadsheets. Data. Now.

BREAKING
Productivity Hacks

Waymo’s Approach to AI Evaluation and Safety

Waymo's Approach to AI Evaluation and Safety - ai safety
Waymo’s Approach to AI Evaluation and Safety

Few companies face higher stakes when deploying AI than Waymo, the self-driving car company under Alphabet that spun out of Google. Its models do not merely generate text or automate back-office tasks: They help vehicles respond to human drivers and make split-second decisions in the physical world. Manasi Joshi, Waymo’s director of engineering for systems intelligence and machine learning, has played a key role in developing the company’s approach to AI evaluation and safety.

The methods Waymo uses to manage those risks – continuous evaluation, carefully curated data, human oversight and clearly defined business outcomes – offer a broader playbook for enterprises deploying AI agents in nearly any industry. To date, Waymo has driven more than 220 million fully autonomous, or “rider-only,” miles, with 17 times fewer serious crash injuries than human drivers over the same distance, according to the company. Joshi explained at VB Transform 2026 how Waymo trains, tests and deploys AI at scale, highlighting the importance of evaluation in the development process.

Table of Contents

Autonomous Vehicle Safety Challenges

High-stakes AI deployment is a significant challenge for autonomous vehicle companies like Waymo. The company’s models must be able to respond to a wide range of scenarios, from everyday driving situations to rare and unexpected events. Unpredictable street navigation is a major concern, as vehicles must be able to adapt to changing road conditions, pedestrian traffic, and other factors. Split-second decision making is also critical, as vehicles must be able to react quickly and safely in emergency situations. Waymo’s approach to AI evaluation and safety is designed to address these challenges, with a focus on continuous testing and validation.

Waymo’s Evaluation-Centric Development

Waymo has adopted what Joshi called “eval-forced development” or “eval-centric development,” making evaluation a core part of engineering rather than a final check performed before deployment. This approach involves assessing a project’s readiness partly by examining the maturity of the tests surrounding it. The maturity of tests is a key factor in determining whether a project is ready for deployment, and Waymo’s methodology combines datasets, performance metrics and infrastructure capable of operating efficiently at scale. Evaluation is not a one-time task to launch a model, but rather a continuous process spanning driving, simulation and validation. By continuing to evaluate its models after launch, Waymo can ensure that its autonomous vehicles remain safe and effective over time. The company’s eval-centric development approach has allowed it to drive millions of miles with a strong safety record, and its methodology has implications for enterprises building other types of AI applications. For example, companies developing customer service agents or financial systems can apply similar evaluation-centric approaches to ensure the reliability and safety of their systems. According to Joshi, much of Waymo’s quality work has shifted toward evaluations, including tests conducted during model training, after training and inside open-loop and closed-loop simulations. Waymo’s approach to evaluation has been refined over time, with a focus on connecting evaluations to actual business outcomes rather than relying solely on performance metrics. By doing so, organizations can ensure that their AI systems are aligned with their business goals and are operating effectively in real-world scenarios, as seen on the official National Highway Traffic Safety Administration website, https://www.nhtsa.gov.

Related: Bright Machines Hybrid Robot Cell Tackles AI Bottleneck

Continuous Evaluation and Testing

Joshi emphasized that evaluations must continue after launch. This approach is essential for ensuring the safety and reliability of Waymo’s autonomous vehicles. The company’s methodology also involves conducting tests during model training and after training, allowing it to identify potential issues before they become major problems. As a result, much of Waymo’s quality work has shifted toward evaluations, with a focus on continuous testing and validation.

The infrastructure for efficient operation at scale is critical to Waymo’s evaluation process. The company has developed a robust testing framework that allows it to simulate real-world scenarios and evaluate its models in a controlled environment. This framework includes a combination of physical and virtual testing, enabling Waymo to test its vehicles in a wide range of conditions, from sunny weather to heavy rain or snow. With over 220 million fully autonomous miles driven, Waymo’s approach to continuous evaluation and testing has proven to be effective, resulting in 17 times fewer serious crash injuries than human drivers over the same distance.

Importance of Reliable Performance Measurement

Assessing a project’s readiness is a critical step in deploying AI applications, and reliable performance measurement is essential to this process. Joshi explained that Waymo assesses a project’s readiness partly by examining the maturity of the tests surrounding it. This approach has clear implications for enterprises building AI applications, as it highlights the need for reliable measurement of system performance. If a company cannot reliably measure a system’s performance, it may not be ready to place that system into production. For example, a company building a customer service agent may need to evaluate its performance in terms of response accuracy, response time, and user satisfaction. According to Joshi, evaluations must continue after launch, and teams must continue evaluating their systems as underlying models, business processes, user behavior, and incoming data change.

Evaluation Hierarchy and Safety Objectives

Waymo’s approach to AI evaluation is grounded in safety objectives, which are carefully defined and prioritized to ensure the development of reliable and safe autonomous vehicles. The company’s evaluation hierarchy is designed to assess the performance of its AI models in a variety of scenarios, including normal driving conditions, as well as rare and dangerous situations. To achieve this, Waymo relies on first-party driving logs and realistic simulations, which provide a wealth of data on how the AI models perform in different environments. Additionally, Waymo’s evaluation process involves exposure to rare and dangerous scenarios, such as construction zones or unexpected pedestrian behavior, to test the limits of its AI models and ensure they can respond safely and effectively.

The use of realistic simulations is a key aspect of Waymo’s evaluation process, as it allows the company to test its AI models in a controlled and safe environment. These simulations can be used to recreate a wide range of scenarios, from everyday driving situations to rare and dangerous events, and can be repeated multiple times to gather detailed data on the performance of the AI models.

Testing Rare and Dangerous Cases

Model-Quality Measurements and Trustworthiness

Waymo’s approach to evaluating its autonomous vehicle systems relies heavily on the trustworthiness of model-quality measurements. The company’s evaluation process involves assessing the performance of its models on a wide range of datasets, including those that reflect real-world driving scenarios. The properties of these datasets are critical, as they must be representative of the various conditions that the autonomous vehicles will encounter on the road. Evaluation data behind performance claims is also essential, as it provides a clear understanding of how the models are performing and where they may need improvement.

Related: Retro Rabbit SmarTek21 Launch SA UX Design Hub

The evaluation process at Waymo is designed to provide a thorough understanding of the company’s autonomous vehicle systems. By using a combination of datasets, performance metrics, and infrastructure, Waymo is able to assess the performance of its models and identify areas for improvement. This approach has enabled the company to develop autonomous vehicle systems that are safe and reliable, with a strong track record of performance on public roads. As the company continues to develop and refine its autonomous vehicle technology, the importance of trustworthy model-quality measurements will only continue to grow.

Waymo’s History and Milestones

With over 220 million fully autonomous miles driven, Waymo has established itself as a leader in the development of autonomous vehicle technology. The company’s history is marked by a series of significant milestones, including the adoption of an eval-centric development approach. Some of the key milestones in Waymo’s history include:

  • 2009: Waymo begins development of its autonomous vehicle technology, with a focus on creating a safe and reliable system.
  • 2012: The company completes its first fully autonomous drive on public roads, marking a major milestone in the development of its technology.
  • 2015: Waymo adopts an eval-centric development approach, which emphasizes the importance of continuous evaluation and testing in the development process.
  • 2018: The company launches its Waymo One service, which provides fully autonomous rides to the public.
  • 2020: Waymo reaches a major milestone, with over 100 million fully autonomous miles driven.
  • 2022: The company expands its Waymo One service to new cities, marking a significant expansion of its autonomous vehicle technology.

Today, Waymo is a subsidiary of Alphabet, and its technology has been developed in collaboration with a range of partners, including the National Highway Traffic Safety Administration. With its strong track record of performance and its commitment to safety, Waymo is well-positioned to continue leading the development of autonomous vehicle technology in the years to come.

Comparison to Other AI Applications

Enterprises building customer service agents, coding assistants, financial systems or other AI applications can learn from Waymo’s approach to evaluation and safety. The implications are significant for customer service agents, as they interact with customers and must provide accurate and helpful responses. Similarly, coding assistants must be able to provide reliable and accurate code suggestions to avoid introducing errors. Financial systems, which handle sensitive financial information, must be extremely reliable and secure. The importance of continuous evaluation and testing helps to ensure that these systems are functioning as intended and can adapt to changing circumstances. For example, a customer service agent may need to be retrained on new products or services, and its performance must be reevaluated to ensure it can provide accurate information.

Other AI applications, such as those used in financial systems, also require careful evaluation and testing. These systems must be able to detect and prevent fraud, and ensure the security and integrity of financial transactions. This is particularly important in industries where the stakes are high, such as finance or healthcare, where errors or security breaches can have serious consequences. According to the official website of the National Institute of Standards and Technology, https://www.nist.gov, security and reliability are essential for building trust in AI systems.

Related: BERT Tests Simplify Proof of Network Success

Practical Advice for Enterprises

For enterprises, this means that testing an agent before launch is insufficient. Those evaluations should also connect to actual business outcomes, rather than relying solely on metrics such as accuracy or precision. For instance, a company deploying a chatbot to handle customer inquiries should continually evaluate its performance and adjust its training data and algorithms as needed to improve customer satisfaction and reduce support costs.

Testing rare and dangerous cases is also essential, as Waymo’s experience has shown. This requires careful consideration of potential scenarios and the development of targeted tests to ensure that the AI system can handle them. This approach can help build trust in AI systems and ensure that they are used to improve business outcomes, rather than simply automating existing processes. The key is to make evaluation a core part of the development process, rather than an afterthought, and to continually assess and improve the performance of AI systems over time. As a result, organizations can create AI systems that are both reliable and effective, and that provide real value to their customers and users.

Impact on the Autonomous Vehicle Industry

Waymo’s approach to AI evaluation and safety has significant implications for the autonomous vehicle industry. By making evaluation a core part of engineering, rather than a final check performed before deployment, Waymo has achieved impressive results, with 17 times fewer serious crash injuries than human drivers over the same distance. This approach has the potential to be adopted by other companies in the industry, leading to a safer and more reliable autonomous vehicle experience. The industry as a whole can benefit from Waymo’s methodology, which combines datasets, performance metrics, and infrastructure capable of operating efficiently at scale. As the industry continues to evolve, it is likely that we will see a shift towards more rigorous evaluation and testing of autonomous vehicles, leading to increased public trust and adoption.

The future of autonomous vehicle development is closely tied to the development of more advanced AI evaluation and safety protocols. As autonomous vehicles become more prevalent on the roads, the need for reliable and efficient evaluation methods will only continue to grow. Waymo’s approach has set a high standard for the industry, and other companies will need to follow suit in order to ensure public safety and gain regulatory approval. With the autonomous vehicle market expected to continue to expand in the coming years, the importance of effective AI evaluation and safety protocols will only continue to increase.

Advances in AI evaluation and safety are likely to have a significant impact on the development of autonomous vehicles in the coming years. As AI technology continues to evolve, we can expect to see more sophisticated evaluation methods and safety protocols emerge. One potential area of development is the use of simulation-based testing, which allows for the evaluation of autonomous vehicles in a controlled and repeatable environment. This approach has the potential to reduce the time and cost associated with traditional testing methods, while also improving the overall safety and reliability of autonomous vehicles. Additionally, the development of more advanced sensor systems and machine learning algorithms will also play a critical role in the future of autonomous vehicle development. Joshi’s work at Waymo has demonstrated the importance of continuous evaluation and testing in the development of autonomous vehicles, and it is likely that this approach will become increasingly prevalent in the industry as a whole. The use of autonomous vehicles is expected to continue to grow, with many experts predicting that they will become a common sight on roads around the world in the near future.

Questions Readers Often Ask

How does Waymo evaluate the safety of its AI systems?

Waymo evaluates the safety of its AI systems through a combination of simulation testing, real-world testing, and rigorous validation protocols. This approach allows the company to identify and address potential safety risks before deploying its autonomous vehicles on public roads. By using a multi-faceted evaluation process, Waymo can ensure its AI systems meet the highest safety standards.

What role does simulation play in Waymo’s AI evaluation process?

Simulation plays a critical role in Waymo’s AI evaluation process, allowing the company to test its autonomous vehicles in a wide range of scenarios and environments. Simulation testing enables Waymo to evaluate its AI systems in a controlled and repeatable manner, which helps to identify potential safety issues and improve overall system performance. This approach also reduces the need for physical prototype testing, making the development process more efficient.

How does Waymo ensure its AI systems can handle edge cases and unexpected events?

Waymo ensures its AI systems can handle edge cases and unexpected events by testing them in simulated scenarios that mimic real-world conditions. The company’s simulation platform generates a wide range of scenarios, including unusual or unexpected events, to evaluate the AI system’s response and decision-making capabilities. By testing its AI systems in these scenarios, Waymo can identify areas for improvement and refine its systems to better handle unexpected events.

What safety protocols does Waymo have in place for its autonomous vehicles?

Waymo has a range of safety protocols in place for its autonomous vehicles, including multiple redundancies and fail-safes to ensure the vehicle can safely respond to any situation. The company’s autonomous vehicles are also equipped with a range of sensors and mapping technologies that provide a 360-degree view of the environment, enabling the AI system to make informed decisions. Additionally, Waymo’s vehicles are designed to be highly reliable and fault-tolerant, with built-in mechanisms to prevent or mitigate potential safety risks.

How does Waymo validate the performance of its AI systems in real-world environments?

Waymo validates the performance of its AI systems in real-world environments through a combination of on-road testing and data analysis. The company’s autonomous vehicles are equipped with a range of sensors and data logging equipment, which provides detailed insights into the AI system’s performance in various scenarios and conditions. By analyzing this data, Waymo can refine its AI systems and improve their performance in real-world environments.

What is Waymo’s approach to addressing potential biases in its AI systems?

Waymo’s approach to addressing potential biases in its AI systems involves rigorous testing and validation to ensure the systems are fair, transparent, and unbiased. The company uses a range of techniques, including data analysis and simulation testing, to identify and mitigate potential biases in its AI systems. By prioritizing fairness and transparency, Waymo can develop AI systems that are trustworthy and reliable.

How does Waymo collaborate with regulatory agencies to ensure its AI systems meet safety standards?

Waymo collaborates with regulatory agencies to ensure its AI systems meet safety standards by providing detailed information about its autonomous vehicle technology and safety protocols. The company works closely with regulators to understand their requirements and concerns, and provides insights into its testing and validation processes. Through this collaborative approach, Waymo can help shape the development of regulatory frameworks that support the safe deployment of autonomous vehicles.

Tags:

Leave a Reply

Your email address will not be published. Required fields are marked *