How to test an LLM feature: output, cost, speed and prompt injection

Engineer typing on a laptop keyboard
Testing a feature built on a language model? Check four things: the output, the cost per answer, the speed, and whether prompt injection can hijack it.
Back to blog

How to test an LLM feature: output, cost, speed and prompt injection

A lot of products now ship at least one feature built on a large language model: a support assistant, a summary button, a search box that answers in sentences. Teams that use these models heavily in development often ask us the same thing when they look for QA. Do we need a very particular kind of tester for this, or does ordinary testing still apply?

Mostly, ordinary testing still applies, with four checks added on top. You test whether the output is what you expected, what each answer costs, how fast it arrives, and whether someone can talk the feature into doing something it should not. This article walks through each one.

Article details
Scope
LLM features inside a product, not model training
Audience
Product owners, engineering leads, QA leads
Reference
OWASP Top 10 for LLM Applications
Clutch rating
4.9 / 5 from 65+ reviews

Why an LLM feature breaks the usual expected result

Most software testing rests on a simple contract: given this input, the system returns this output, every time. A language model does not keep that contract. Ask the same question twice and you can get two different, equally reasonable answers. Change a word in the prompt your developers wrote and the tone, length or accuracy of every answer can shift.

That does not make the feature untestable. It changes what you assert. Instead of checking that the answer matches a fixed string, you check properties of the answer: does it contain the facts it must contain, does it stay inside the topic, does it avoid the things it must never say, does it arrive in the format the rest of the product expects. A tester who has spent years writing precise expected results needs a short adjustment here, not a new career.

Note the boundary of this article. If you are training or fine-tuning your own model, you also need data quality, bias and drift testing, which we cover in how to test AI systems. Here we are talking about the far more common case: a product feature that calls a model someone else trained.

Output: is the answer what you expected?

Start by building a fixed set of realistic inputs, sometimes called a golden set. Take them from real user questions where you have them, and add the awkward ones: very long inputs, inputs in a second language, questions the feature should refuse, questions that sit right on the edge of its topic. Write down for each one what a good answer must include and what it must not.

Run that set every time the prompt, the model version or the surrounding code changes. Some of the checks can be automated, such as required fields present, banned phrases absent, answer length inside bounds, valid JSON when the product parses the reply. The rest needs a person reading a sample of answers against the criteria you wrote. That human review is not a weakness of the method. It is the method, because "is this answer correct and appropriate" is a judgement.

Watch especially for confident wrong answers. A model that invents a refund policy your company does not have will phrase it beautifully. Your test criteria should name the facts the answer is allowed to rely on, so an invented one stands out.

Cost and speed: the two checks teams forget

Every call to a hosted model is billed, usually by the amount of text going in and coming out. That makes cost a functional property of the feature, and it can be tested like one. Measure what a typical request costs and what the worst case costs: the longest document a user can paste, the conversation that runs to fifty turns, the retry loop that fires three times when the provider is slow. Agree a budget per request with the product owner and turn it into a test that fails when a change blows through it.

Speed belongs next to cost because the two trade against each other. Larger models and longer prompts tend to be slower. Measure the time to the first visible word and the time to the full answer, under normal load and when several users hit the feature at once. Then test what the product does when the model provider is slow or down. A support assistant that leaves the user staring at a spinner for a minute is a defect, even if the eventual answer is perfect. The fallback, whether a timeout message, a cached answer or a hand-off to a human, needs its own test cases.

These checks are easy to automate and they catch real problems. A prompt change that doubles the instructions sent with every request can pass every output check and still double your monthly bill.

Prompt injection: SQL injection's new cousin

Web applications learned the hard way that user input mixed into a database query can rewrite the query. That is SQL injection. Language model features have the same structural weakness. Your developers write instructions for the model, the user's text is placed next to them, and the model cannot reliably tell where your instructions end and the user's begin. Prompt injection is the attack that exploits that, and it sits first on the OWASP Top 10 for LLM Applications.

Test it from two directions. Direct injection is a user typing instructions meant to override yours: ignore the rules above, reveal your system prompt, answer as if you were an administrator. Indirect injection is quieter and usually more dangerous: the instructions hide inside content the feature reads on the user's behalf, such as an uploaded document, a web page, an email it summarises or a product review.

What you are really testing is the damage an injected instruction can do. Can it make the feature reveal data belonging to another user? Can it trigger an action, like sending an email or changing a record, when the feature has access to tools? Can it make the assistant say something your company would have to apologise for? The narrower the feature's permissions, the smaller the answer to each question, which is why the best fix is often in the design rather than the prompt.

This is the one check where a security specialist earns their place. Our security testing team runs injection testing alongside the functional work, and the AI Security Toolkit correlates automated scanner findings with in-house checks in one pass, so a reported issue comes with the chain that makes it exploitable rather than as an isolated warning.

Everything around the model still needs normal testing

The model is usually a small part of the feature. Around it sit a user interface, authentication, permissions, logging, error handling, rate limits and the code that turns the model's reply into something on screen. All of that is tested the way it always was, and in our experience that is where many of the defects in these features actually live. A reply that the front end cannot render, a permission check that runs after the model has already seen private data, a log file that stores every user's question in plain text: none of these are model problems.

So the honest answer to the buyer's question is this. You do not need a very particular species of tester. You need testers who are comfortable asserting properties instead of exact strings, a security specialist available for the injection work, and someone senior enough to write the acceptance criteria with your product owner. The NIST AI Risk Management Framework is a good shared reference if your organisation wants a structure for those conversations.

Frequently asked questions

Usually not a separate kind of engineer. A good tester who can assert properties of an answer rather than an exact string covers output, cost and speed. Prompt injection is the part where a security specialist should be involved.
Test properties instead of exact text. Keep a fixed set of realistic inputs, write down what each answer must and must not contain, automate the checks that can be automated, and have a person review a sample against the written criteria after every change.
It checks whether text supplied by a user, or hidden in content the feature reads, can override the instructions the developers gave the model. The goal is to find out what an injected instruction could do: leak data, trigger actions, or produce harmful replies.
Yes. Each call is billed by the amount of text in and out, so a prompt change can pass every quality check and still double the running cost. Agree a budget per request and add a test that fails when a change exceeds it.

Shipping a feature built on a language model?

We test the feature and everything around it: output, cost, speed, injection and the ordinary code that holds it together. Start with two weeks on your own product.

Talk to our team

Need help with software testing?

BetterQA provides independent QA services across manual testing, automation, security audits, and performance testing. ISO 27001, 9001, 14001 and 13485 certified.

Share the Post: