How to test an LLM feature: output, cost, speed and prompt injection
A lot of products now ship at least one feature built on a large language model: a support assistant, a summary button, a search box that answers in sentences. Teams that use these models heavily in development often ask us the same thing when they look for QA. Do we need a very particular kind of tester for this, or does ordinary testing still apply?
Mostly, ordinary testing still applies, with four checks added on top. You test whether the output is what you expected, what each answer costs, how fast it arrives, and whether someone can talk the feature into doing something it should not. This article walks through each one.
Why an LLM feature breaks the usual expected result
Most software testing rests on a simple contract: given this input, the system returns this output, every time. A language model does not keep that contract. Ask the same question twice and you can get two different, equally reasonable answers. Change a word in the prompt your developers wrote and the tone, length or accuracy of every answer can shift.
That does not make the feature untestable. It changes what you assert. Instead of checking that the answer matches a fixed string, you check properties of the answer: does it contain the facts it must contain, does it stay inside the topic, does it avoid the things it must never say, does it arrive in the format the rest of the product expects. A tester who has spent years writing precise expected results needs a short adjustment here, not a new career.
Note the boundary of this article. If you are training or fine-tuning your own model, you also need data quality, bias and drift testing, which we cover in how to test AI systems. Here we are talking about the far more common case: a product feature that calls a model someone else trained.
Output: is the answer what you expected?
Start by building a fixed set of realistic inputs, sometimes called a golden set. Take them from real user questions where you have them, and add the awkward ones: very long inputs, inputs in a second language, questions the feature should refuse, questions that sit right on the edge of its topic. Write down for each one what a good answer must include and what it must not.
Run that set every time the prompt, the model version or the surrounding code changes. Some of the checks can be automated, such as required fields present, banned phrases absent, answer length inside bounds, valid JSON when the product parses the reply. The rest needs a person reading a sample of answers against the criteria you wrote. That human review is not a weakness of the method. It is the method, because "is this answer correct and appropriate" is a judgement.
Watch especially for confident wrong answers. A model that invents a refund policy your company does not have will phrase it beautifully. Your test criteria should name the facts the answer is allowed to rely on, so an invented one stands out.
Cost and speed: the two checks teams forget
Every call to a hosted model is billed, usually by the amount of text going in and coming out. That makes cost a functional property of the feature, and it can be tested like one. Measure what a typical request costs and what the worst case costs: the longest document a user can paste, the conversation that runs to fifty turns, the retry loop that fires three times when the provider is slow. Agree a budget per request with the product owner and turn it into a test that fails when a change blows through it.
Speed belongs next to cost because the two trade against each other. Larger models and longer prompts tend to be slower. Measure the time to the first visible word and the time to the full answer, under normal load and when several users hit the feature at once. Then test what the product does when the model provider is slow or down. A support assistant that leaves the user staring at a spinner for a minute is a defect, even if the eventual answer is perfect. The fallback, whether a timeout message, a cached answer or a hand-off to a human, needs its own test cases.
These checks are easy to automate and they catch real problems. A prompt change that doubles the instructions sent with every request can pass every output check and still double your monthly bill.
Prompt injection: SQL injection's new cousin
Web applications learned the hard way that user input mixed into a database query can rewrite the query. That is SQL injection. Language model features have the same structural weakness. Your developers write instructions for the model, the user's text is placed next to them, and the model cannot reliably tell where your instructions end and the user's begin. Prompt injection is the attack that exploits that, and it sits first on the OWASP Top 10 for LLM Applications.
Test it from two directions. Direct injection is a user typing instructions meant to override yours: ignore the rules above, reveal your system prompt, answer as if you were an administrator. Indirect injection is quieter and usually more dangerous: the instructions hide inside content the feature reads on the user's behalf, such as an uploaded document, a web page, an email it summarises or a product review.
What you are really testing is the damage an injected instruction can do. Can it make the feature reveal data belonging to another user? Can it trigger an action, like sending an email or changing a record, when the feature has access to tools? Can it make the assistant say something your company would have to apologise for? The narrower the feature's permissions, the smaller the answer to each question, which is why the best fix is often in the design rather than the prompt.
This is the one check where a security specialist earns their place. Our security testing team runs injection testing alongside the functional work, and the AI Security Toolkit correlates automated scanner findings with in-house checks in one pass, so a reported issue comes with the chain that makes it exploitable rather than as an isolated warning.
Everything around the model still needs normal testing
The model is usually a small part of the feature. Around it sit a user interface, authentication, permissions, logging, error handling, rate limits and the code that turns the model's reply into something on screen. All of that is tested the way it always was, and in our experience that is where many of the defects in these features actually live. A reply that the front end cannot render, a permission check that runs after the model has already seen private data, a log file that stores every user's question in plain text: none of these are model problems.
So the honest answer to the buyer's question is this. You do not need a very particular species of tester. You need testers who are comfortable asserting properties instead of exact strings, a security specialist available for the injection work, and someone senior enough to write the acceptance criteria with your product owner. The NIST AI Risk Management Framework is a good shared reference if your organisation wants a structure for those conversations.
Frequently asked questions
Shipping a feature built on a language model?
We test the feature and everything around it: output, cost, speed, injection and the ordinary code that holds it together. Start with two weeks on your own product.
Talk to our teamNeed help with software testing?
BetterQA provides independent QA services across manual testing, automation, security audits, and performance testing. ISO 27001, 9001, 14001 and 13485 certified.