How to test an integration when the partner’s test environment doesn’t match production

BetterQA mug on a meeting table while the team works
The partner's sandbox does not behave like production. How to test your side, simulate the provider, list the differences and check safely after release.
Back to blog

How to test an integration when the partner's test environment doesn't match production

Most products lean on someone else's system for the part that matters most. A payment provider takes the money, a bank confirms the transfer, a shipping carrier books the parcel, an identity provider decides who gets in. Many of those providers offer a sandbox or UAT environment. Then the team discovers that the sandbox answers differently from production, or the credentials never arrived, or the test environment cannot reach the provider's servers at all.

At that point we often hear the same conclusion: the riskiest part of the product is untestable. It is not. Some of the risk cannot be removed before release, but most of it can be tested, and the rest can be reduced and watched. This article explains how.

Article details
Question
How do I test an integration when the partner's sandbox does not behave like production?
Audience
CTOs, product owners and founders whose product depends on third-party providers
BetterQA team
50+ engineers, founded 2018
Clutch rating
4.9 / 5 from 65+ reviews

What can you test when you cannot reach the provider?

An integration has two halves. The provider owns theirs. You own yours: the requests you send, and everything your product does with the answer. Most production incidents we see around third-party systems sit on your half. The provider replied with something legitimate, and the product did not know what to do with it.

That half is fully testable without the provider. Write down every request your product sends and every answer it can get back, then test how the product behaves for each one. A useful starting list:

  • The request itself. Correct fields, formats, currency and amount handling, character encoding in names and addresses, and the authentication headers.
  • Every documented response. Success, each decline or rejection code, validation errors, and the "pending" or "under review" states that resolve later.
  • Timeouts. What the user sees when the provider takes thirty seconds, and what state the order or account is left in.
  • Retries. Whether a retry can charge a customer twice or book two shipments. Ask whether the provider supports idempotency keys and whether you send them.
  • Duplicates and late arrivals. The same callback delivered twice, or a webhook that arrives before the API response it belongs to.
  • Outages. A 500, a dropped connection, an expired certificate. The product should fail in a way a person can recover from.

None of this needs the provider to be reachable. Our integration testing guide covers the wider technique.

How do you build a stand-in that behaves like production?

Once you know what to test, you need something to answer your requests. That is a mock, a stub, or at larger scale service virtualisation: a fake provider that returns the responses you choose, including the awkward ones a real sandbox will never produce on demand.

The weakness of any simulator is that someone wrote it, and it encodes what they assumed. A mock built from the provider's documentation tests your product against the documentation. Production does not always agree with the documentation. So build the stand-in from real production behaviour wherever you can:

  • Capture real responses from production logs, with personal and payment data removed, and replay them as fixtures.
  • Include the odd ones: the extra field nobody documented, the error message in a different language, the empty string where you expected a null.
  • Add contract tests that check the shape of what you send and what you expect back. When the provider changes a field, a contract test is the first place it shows up.

Then keep the recordings current. A fixture captured two years ago describes a provider that may no longer exist. Treat refreshing them as part of maintenance, the same way you treat an automated test suite that needs updating when the interface changes.

Why should you write the sandbox differences down?

Teams usually know their sandbox is different. What they rarely have is a written list of how. Without one, every test passed in the sandbox carries a silent caveat nobody can read, and the release decision is made on a feeling.

Keep a short table with one row per known difference: what the sandbox does, what production does, and how you cover the gap. Typical rows look like this. The sandbox approves every card, production declines some for fraud reasons. The sandbox replies instantly, production sometimes takes several seconds. The sandbox never sends a duplicate webhook, production occasionally does. The sandbox has looser rate limits. The sandbox certificate or endpoint differs from the live one.

For each row the answer is one of three things: tested in the simulator, checked in production after release, or accepted as a known risk with someone's name next to it. That table is worth more than the sandbox test results, because it tells the person signing off exactly what has and has not been proven. It belongs in the release checklist, which our article on QA's role in release management describes.

What can you safely check in production?

Some things only production can tell you. Better to plan for that than to pretend the sandbox settled it. Testing in production does not mean letting customers find the bugs. It means small, deliberate checks with a way back.

  1. Put the integration behind a feature flag. You can switch it off in minutes without a deploy if the provider starts behaving badly.
  2. Roll out in stages. Internal users first, then a small share of traffic, then everyone. Each stage is a test with real answers from the real provider.
  3. Run limited live checks where it is safe. A small real transaction with a company card that is refunded the same day, a test shipment to your own office, a login with a dedicated test identity. Agree these with the provider and with whoever owns compliance first.
  4. Watch the numbers that move when it breaks. Decline rate, timeout rate, callbacks received versus requests sent, and orders stuck in a pending state. Alert on a change, not just on an error.

Logging is what turns a production failure from a mystery into a ticket. For every call to the provider, record the request identifier, the idempotency key, the timestamp, the response code, how long it took and which retry it was. Never log card numbers, tokens or passwords. When something goes wrong at night, this is what lets someone answer "did we send it, did they receive it, and what did they say" before the morning.

What risk remains, and where should a new QA vendor start?

After all of the above, some risk is still there. You cannot prove before release that a provider's production fraud rules will treat your customers the way you expect, or that their system will behave under the load of your launch day. You can reduce that risk, detect it quickly and limit how many people it reaches. A test report that says otherwise is not being straight with you. A good one says which parts were proven in test, which will be confirmed in production, and which are accepted.

One practical point if you are bringing in a QA vendor for the first time. Do not make the integration nobody can reach their first assignment. Pick a first area where the testers can reach the backend, see real responses and file bugs a developer can confirm the same day. The hard integration comes second, when they know your system well enough to build a sensible simulator and write the differences table with you. Our article on what a QA trial should prove goes into choosing that first area.

This is also work we do as part of a managed testing engagement: mapping your integrations, building the stand-ins in your repository and CI, and agreeing the production checks with your team. What we build belongs to you. Engagements run at $25 to $45 an hour, and a team is usually working on your product about two weeks after the first call.

Frequently asked questions

Sandbox testing runs against the provider's test environment, which is safe but often simplified: it may approve everything, reply instantly and never send duplicates. Production testing means small, controlled checks against the live provider after release, behind a feature flag and with monitoring. Most integrations need both, plus a written list of how the two differ.
Yes, most of it. Test your own side: the requests you send and how the product handles every response, timeout, retry, duplicate and outage. Use mocks or service virtualisation built from real production responses, and add contract tests so changes on the provider's side show up early.
It can be, if the checks are small and reversible. Use a feature flag, roll out in stages, run limited live transactions that you can refund or cancel, and agree them with the provider first. Monitor the rates that change when the integration breaks, and log enough to diagnose a failure without storing sensitive data.
No. Some behaviour only exists in the provider's production system. The goal is to prove what can be proven in test, confirm the rest quickly after release, and make the remaining risk visible and owned rather than silent.

Not sure how much of your integration is really tested?

Tell us which provider worries you. We will map what can be tested before release, what needs a production check, and what you are currently accepting without knowing it.

Book a call See our testing services

Need help with software testing?

BetterQA provides independent QA services across manual testing, automation, security audits, and performance testing. ISO 27001, 9001, 14001 and 13485 certified.

Share the Post: