Using AI testing tools when your data can't leave the company
Plenty of companies cannot connect their Jira, their repositories or their screen recordings to a model hosted by someone else. The reason is usually simple: client data sits in those places. Meanwhile almost every AI-assisted testing tool works by sending something to a model, whether it is generating test cases from requirements, writing bug reports from recordings or scanning code.
That does not mean these tools are off the table. It means the evaluation has to start with where your data goes, not with the demo. This is the checklist we would use if we were on your side of the table, written for the CTO, product owner or founder who has to say yes or no and then defend it.
Which parts of your testing actually send data to a model?
Start by listing the inputs, because "the tool uses AI" is too vague to approve or reject. Ask the vendor to map every feature to the data it sends. If they cannot do that in writing, you have your first answer.
- Tickets and requirements. Test case generation reads your Jira stories, acceptance criteria and comments. Those often quote customers, paste in support emails or name the client a feature was built for.
- Source code. Code scanning and test generation from code send files, sometimes whole repositories. Code can contain hard-coded keys, internal hostnames and comments nobody meant to share.
- HAR files and logs. These are the most underestimated. A browser HAR file records every request and response, including cookies, session tokens and authorisation headers. Application logs routinely carry email addresses, IDs and occasionally passwords.
- Screen recordings and screenshots. A recording of a bug on a staging environment loaded with copied production data shows real names, balances and addresses on every frame.
- Test data. If your test environment was seeded from production, every test that runs against it handles real records, whether or not a model is involved.
Once the list exists, you can decide surface by surface. Many teams find they are comfortable sending a sanitised requirement to a hosted model and not comfortable sending a HAR file under any circumstances.
Does the vendor train on your data, and where does the model run?
These are two separate questions, and vendors sometimes answer the first one and let you assume the second.
Training
Ask directly whether anything you send is used to train or improve a model, theirs or their model provider's. Then ask where that promise is written. A sentence on a marketing page is not the same as a clause in your contract or the model provider's terms for business customers. Ask, too, how long they keep what you send, and whether people at the vendor or provider can read it.
Hosting
"Where does the model run" has at least four very different answers, and they are different builds with different owners of the cost:
- Vendor-hosted. The tool calls a model the vendor or its provider runs. Simplest for you, and the case your constraint usually rules out.
- A cloud model inside your own tenant. The model runs in your cloud account, under your agreements, and data stays inside your boundary. You pay the cloud bill and someone has to set it up and maintain access.
- A self-hosted open model. You run an open-weight model on infrastructure you control. More control, but output quality varies by model, and someone owns updates, scaling and security patching.
- On-premise hardware. The model runs on machines in your building. Full control, and the highest cost in hardware, power and people.
When a vendor says "we can run it privately", ask which of these they mean, who pays for it, and whether the features you saw in the demo work the same way in that setup.
Does the rule apply to a proof of concept on synthetic data?
Teams often stall here: engineering wants a tool, security says no. A useful question to put to your security or compliance owner is whether the restriction is about the tool or about the data. If it is about the data, a proof of concept on synthetic data, a demo project or a public open-source repository may be perfectly acceptable.
Be clear about what that proves. A tool that writes good test cases from a tidy synthetic requirement may do worse on your real backlog, with its abbreviations, half-finished stories and internal jargon. And synthetic data says nothing about data handling, which is the part you need approved.
Who decides
The decision belongs to whoever owns information security or compliance in your company, not to the engineering lead who found the tool and not to the vendor. Bring them in early, give them the surface-by-surface list from above, and let them say which data classes can go where. In companies with ISO 27001 or similar programmes there is usually already a supplier assessment process. The NIST AI Risk Management Framework is a neutral reference if your team wants a shared vocabulary for the discussion.
What should you scrub, and who checks the output?
Even with a model that never leaves your tenant, two habits matter more than the hosting choice.
Scrub secrets from logs and HAR files first
Make it a rule that nothing goes into a ticket, a tool or a model until secrets are removed. That means stripping cookies, bearer tokens, authorisation headers and API keys from HAR files, and masking emails, account numbers and IDs in logs. Do it with a script, not by eye. The OWASP Logging Cheat Sheet lists the kinds of data that should never reach a log in the first place, which is the better fix. This applies to human testers too.
A human reviews every generated artefact
Generated test cases and bug reports are drafts. A model can invent a requirement that does not exist, miss the edge case that matters, or copy a piece of sensitive data from its input into its output. Someone who knows the product has to read each one before it counts as a test or a defect. Ask any vendor who does that review. If output goes straight into your Jira, the checking work is yours. Testing generated output is its own discipline, which we cover in our guide on how to test features built on language models.
What should the contract say?
Before signing with an AI testing tool or a QA vendor that uses one, check that the agreement covers these points:
- Your data is not used to train or improve any model, and the same applies to every sub-processor.
- A named list of sub-processors, including the model provider, with notice before it changes.
- Where data is processed and stored, and how long inputs and outputs are kept.
- What happens to your data when the contract ends, and how deletion is confirmed.
- Who owns what the vendor produces: test cases, scripts, reports.
If you are buying testing as a service rather than a tool, add one more question: does the vendor work in your tools, or does your data have to move into theirs? A team that files bugs in your Jira and commits to your repository leaves far less data outside your boundary than one that needs exports. If the work includes security testing, the same questions apply, and our security testing service page explains what that work covers.
Where BetterQA stands on these questions
We build and use our own testing tools, and we get asked these questions often, so here are our answers plainly. No AI training happens on our platforms, on anyone's data. Every test case and bug report our tools generate goes through a human engineer before it counts. We work in the client's tools: your Jira, your repository, your CI. What we build belongs to you. BetterQA is ISO 9001 and ISO 27001 certified.
Some answers depend on your setup and what your security team allows, so we would rather go through them with you than guess. Bring the checklist to a call, ask us each question, and judge the answers the same way you would judge any vendor. If you are still deciding how much of your testing to hand over, our managed testing services page explains how we run the function and own the result.
Frequently asked questions
Ask us the hard questions on a call
Bring this checklist and your security owner. We will answer each question for how we would work on your product, and tell you plainly where something depends on your setup.
Book a call See our servicesNeed help with software testing?
BetterQA provides independent QA services across manual testing, automation, security audits, and performance testing. ISO 27001, 9001, 14001 and 13485 certified.