The Global AI Assurance Sandbox is an international initiative by IMDA and the AI Verify Foundation to create a testing ground for builders or deployers of GenAI applications to get them tested by specialist technical testers. Building on the success of the pilot phase, the Sandbox is an ongoing initiative where AI deployers and technical testing firms can participate when ready.
01
02
03
30
12
14
















You are launching or deploying a GenAI or Agentic AI app, and are looking for:
GenAI applications:
You are offering AI testing as a service/product, and are looking for opportunities to:
Begin by using IMDA’s Starter Kit to identify what and how to test. The Starter Kit supports testing in areas such as:
You are also encouraged to explore other use case–specific considerations, including:
By participating in Sandbox, you will produce a case study report and be featured alongside others on our website and relevant events
Submit interest to join the Sandbox
Your submissions will be assessed
AI deployers and testers will be introduced to each other based on objectives, scope and preference, and you decide on your partner based on commercial discussions
Collaborate to complete the technical testing and produce case study report within 3 months
Your case study report is submitted to IMDA/AIVF and published
How meticulous that companies need to be in not only defining the bad behavior that should not exist in AI applications, but also the good behavior that should exist and what is considered good.
Safeer Mohiuddin
Guardrails AI
We learned that businesses replacing internal processes with LLM tools need outputs that don’t just pass technical accuracy checks, but also precisely match nuanced internal standards. And these standards are often subjective and qualitative, emphasizing engagement, coherence, and suitability for customer-facing communications.
Matthew Dodgson
PwC
AI model testing and AI application testing are different. The definition of safety is very different across different use case and domains. For example, the risk in healthcare is different in the risk in finance.
Yifan Jia
AIDX Tech
Guidelines or best practices on how to build LLM applications so they can be easily tested would be critical. Formal accreditation of vendors and their test approaches could also help in assuring consistency.
Miguel Fernandes
Resaro
There’s too much headache over the cost and complexity of mobilizing testing and assurance technology. Democratize access to these testing technologies well beyond the realm of frontier Labs into the economy, including in SMEs.
Miguel Fernandes
PRISM Eval
Running some tests and computing some numbers, that is the easy part. But knowing what tests to execute and how to interpret the results, that was the hard part.
Dr Martin Saerbeck
AIQURIS