AI Testing: Testing AI Systems Professionally

AI-based systems require different testing strategies than traditional software. Generative AI and large language models, in particular, produce probabilistic results: the same input can lead to different answers, each of which appears plausible.

Formalized test cases with a single, fixed expected result are therefore often insufficient.

Man testing an AI software application

AI Is Changing Software Testing

Traditional testing methods compare an actual result with a clearly defined target result. With AI systems, however, result spaces, probabilities, unexpected behavior, and technical risks must also be taken into account.
Relevant examples include:

Risk-Based Testing of AI

Not every error is equally critical. A linguistically awkward result has a different significance than an incorrect recommendation in a financial, administrative, or healthcare process.

Objentis therefore tailors its testing strategy to the actual risks associated with the application. We don’t just examine whether a system works; we also consider:

in what context it is used,

which individuals are affected by its results,

how serious potential mistakes can be,

how errors are detected and handled,
what controls and escalation procedures are required.
This risk-oriented approach has long been an integral part of Objentis’ economic testing and is becoming even more important in AI testing.

Possible Areas of Examination

Depending on the application, the following quality characteristics, among others, can be examined:

Our Approach to AI Testing

1. Understand the context and intended use
We analyze user groups, processes, data, dependencies, and potential impacts.
2. Identify and Prioritize Risks
Together, we determine which types of misconduct could cause the most damage.
3. Develop a testing strategy
We combine formal verification, exploratory testing, and probabilistic evaluation methods.
4. Examine AI Behavior
The system is confronted with typical, critical, ambiguous, and unexpected inputs.
5. Evaluate the results
Results are evaluated not only from a technical perspective, but also from a subject-matter and risk-related perspective.
6. Develop Action Plans
We identify where technical improvements, controls, monitoring, or human approvals are required.

Expert Experience Remains Crucial

AI can generate test cases and analyze large sets of results. However, it does not decide on its own which risks are acceptable for a company or its customers. Professional AI testing therefore requires people who can think critically, identify connections, and take responsibility for the evaluation.

Would you like to test AI-based software or integrate AI effectively into your quality assurance process?