Starter Kit for Testing LLM-based Applications for Safety and Reliability

Gen AI Apps Bring New Risks and New Ways of Testing

Generative AI apps bring immense opportunities — and equally complex risks. Every prompt and output can have real-world consequences for customers, employees, and society. To deploy AI responsibly, organisations need clarity on what to test, how to test and how to interpret test results.

This is where IMDA’s Starter Kit for Testing LLM-based Applications for Safety and Reliability (“Starter Kit”) for testing comes in.

What is the Starter Kit?​

The Starter Kit is a set of voluntary guidelines for pre-deployment testing of LLM-based applications. It consolidates emerging best practices and methodologies to help organisations assess whether their apps meet baseline safety and reliability needs.

The Starter Kit was developed in consultation with industry – in particular through our Global AI Assurance Sandbox during its Pilot phase – and government to address the new risks from the use of LLMs in apps.

It lays out a structured approach to testing and practical guidance on how to identify relevant risks, test for these risks and assess whether thresholds are met, providing consistency in testing methodology in a rapidly evolving environment.

It tackles five such risks:

Starter Kit Provides Practical Testing Guidance for All Businesses and is Applicable Across Sectors and Use Cases

Step 01: Identify

Determine the relevant risks to test for your app, calibrate the extent of testing required, and define thresholds for baseline safety and reliability.

Step 02: Test

Run tests in a structured manner, from the app’s outputs to its components.

Step 03: Assess

Analyse results and determine whether your safety thresholds (i.e. baseline safety and reliability) have been met, to inform mitigations and next steps.

Who Is This For?

01
Developers and testers whether in-house teams or third parties

Scoping risks and extent of testing and conducting tests in a structured and rigorous manner, in line with industry practices, analysing results

02
Compliance / Responsible AI professionals

Identifying relevant risks, establishing thresholds to assess whether baseline safety and reliability requirements are met

03
Business
leaders

Knowing key risks to look out for, understanding fundamentals of testing 
LLM-based apps

When Should You Use It?

Primarily during the pre-deployment stage when your app is almost ready and you want a final safety or sanity check before launch.

However, you can refer to the Starter Kit throughout the app development lifecycle – e.g. during your design or development to bake in safety principles to support safety (and testing) downstream.

The Starter Kit is grounded in real practices and evolves with industry needs

Updated as technology evolves:

The Starter Kit will continue to expand to address new modalities, emerging risks, and updated testing methodologies as technology develops.

Grounded in real practices:

It will guide the testing of real-life apps in the ongoing Global AI Assurance Sandbox, while translating concrete learnings back into the Starter Kit.

Adapted into technical tools:

The testing guidance in the Starter Kit is adapted into real tools like Project Moonshot, the technical toolkit for GenAI application testing.

The feedback loop between practice, tools, and policy keeps governance agile, grounded, and innovation-friendly. It allows businesses to build trusted GenAI, backed by common standards, open knowledge exchange, and a community committed to safe and responsible innovation.