At least 23 of the 26 apps I audited in June and July 2026 had zero working automated tests, which is the gap software testing services are sold to close. Six kinds of seller answer that search, from a QA outsourcing firm to one freelance tester, and the first question for each is where the test suite lives when the contract ends.

What you are actually buying from software testing services: six sellers, six different things

Software testing services are sold by 6 kinds of seller: a QA outsourcing company, a test automation agency, a QA-as-a-service platform, a managed tester community, a freelance QA engineer and an in-house hire. They hand back different things, from bug reports to a test suite, and the suite does not always end up in your repository.

That figure of at least 23 is out of all 26 apps in my June and July 2026 audits, a selected set of audited apps, not a random sample or a rate for AI-built apps in general.

This is a buying guide inside the wider job of production hardening: getting an app that already works ready for real users, money and change.

What you buy, underneath the labels, is people and tooling outside your team who check that the software does what it should before your users find out it does not. In my reading, the kinds of testing most often sold under the phrase are four: manual testing, test automation, performance testing and security testing.

A firm that calls itself a software testing services company usually sells row 1 or row 2 below, in my reading, and the difference matters more than the name: row 1 rents you testers, row 2 builds you code. Whatever a software test company calls its offer, the table sorts it by what comes back to you. The “what they do”, “sized for” and “who owns the suite” cells are my reading of each seller type from its own published pages; the billing cells come from the sellers’ own sites.

SellerWhat they doWhat you get backHow it billsSized forWho owns the suite afterwards
1. QA outsourcing or software testing companyA managed team of manual and automation testers runs your test cyclesTest plans, bug reports, a regression packEngagement models the firm names, among them a dedicated team, fixed price and time and material (QA Mentor’s list)A product team with a release cadenceWhatever the contract says, so ask before signing
2. Test automation agencyEngineers build a test framework and scripts, then maintain themA test suite plus a maintenance arrangementNot stated on the seller pages analyzed for this searchA team that already has manual test cases to automateYou, if the suite sits in your repository and CI
3. QA-as-a-service platform (QA Wolf, for example)A self-serve platform, or a managed service that writes and runs testsTests run on the vendor’s infrastructure; on the managed service, also written and maintained for youUsage-based self-serve Platform, or managed Coverage as a ServiceA team that wants coverage without hiringThe vendor runs them; exportable, by QA Wolf’s own statement (quoted further down)
4. Managed tester community (Testlio, for example)Testers matched to your app through a platformBug reports and test run resultsA platform subscription plus a consumption fundDevice and browser spread at release timeNot stated on Testlio’s pricing page, so ask what you keep if you cancel
5. Freelance QA engineerOne person, by the hour, doing what you directWhatever you asked for: test runs, bug reports, first scriptsHourlyA bounded job with a clear endYou, if the work lands in your repository
6. In-house QA hireThe whole testing job, as a salaried employeeEverything, for as long as they stayA salaryA team shipping every week with a budget for headcountYou

Before comparing the software tester companies in that table, look at what the apps in my audits had to start from. At least 18 of the 21 third-party apps had no working test anywhere: 17 with literally none, plus a retail POS whose checkout “test suite” never executed the actual checkout code. Only 1 of the 21 third-party apps, a healthcare FHIR hub, was credited with a real test suite, and even it skipped sign-up, login and the payment webhook.

Those two counts come from the same June and July audits and cover the 21 third-party apps among the 26; the apps were picked for audit, not drawn at random. So a reader with an app like those starts from nothing, and a seller sized to automate your existing manual test cases is the wrong size for it. That is my reading, and the table is a documented analysis of each seller type’s own published pages, not a hands-on review of any firm.

Software testing companies in the USA, India and everywhere else: what location changes

Distance changes the same things for testing as for any outsourced engineering work, and the published rate bands by region, the hours of overlap and who publishes the ranked lists of outsourcing firms are covered in what changes when the company is in another country.

One point is specific to testing. Testers in any country see whatever data your test environment holds. So a US buyer asks for a staging environment with anonymized data, and for the firm’s data-handling terms in writing, before the first test cycle, wherever the firm sits. Location does not change what you should get back: that is the table above.

What it costs

Software testing prices are hard to find on this search. QA Wolf prints usage rates for its self-serve platform, Testlio describes a platform subscription plus a consumption fund, and none of the seven firm pages analyzed for this search showed a price. A salary sets the in-house floor.

The table gives the shape each seller type publishes, in the seller’s own terms where it prints any. The “what moves it” cells are my reading.

SellerPublished price shapeUnitWhat moves itSource and date
1. QA outsourcing companyNo price printed; billing models listed, among them a dedicated team, fixed price and time and materialNot statedTeam size and how long you keep themSeven seller pages analyzed 2026-09-17; QA Mentor’s home page, 2026-09-27
2. Test automation agencyNo price printedNot statedNumber of journeys, then monthly maintenanceSeven seller pages analyzed 2026-09-17
3. QA-as-a-service platformSelf-serve Platform: “1¢ per AI credit” and “15¢ per runner minute”; Coverage as a Service is custom-pricedAI credits and runner minutes; for the managed service, the number of tests under managementHow many tests run, how often, for how longQA Wolf’s pricing page, 2026-09-27
4. Managed tester communityA platform fee plus an annual consumption fund; no figure printedNot stated on Testlio’s pricing pageTesting volume drawn from the fundTestlio’s pricing page, 2026-09-27
5. Freelance QA engineerHourly; see what the freelance rate bands are built fromPer hourHours, and whether scripts are part of the jobThe linked article
6. In-house QA hireU.S. median annual wage $104,300 (2025), before benefits and toolingPer yearLocation, seniority, benefitsO*NET OnLine, 2026-09-27

The in-house row uses the U.S. figure for software quality assurance analysts and testers, from O*NET’s profile of software quality assurance analysts and testers, which prints “$50.14 hourly, $104,300 annual” as the 2025 median.

Whatever the shape, a quote moves with a few things you can list before any sales call:

  • the number of user journeys to cover (signup, login, checkout and the one action your product exists for)
  • the platforms and devices that matter to your users
  • how often the app ships
  • whether test data and a staging environment already exist
  • maintenance, because tests break when the interface changes

Prices checked 2026-09-27.

When to pick each

The right testing seller depends on two facts about your team: how often you ship, and who will own a failing test. A managed QA team fits a product team on a release cadence. An app with one developer needs a small suite that runs in its own CI before it needs any retainer.

Each option below gets a “fits when” and a “does not fit when”. Both are my reading, not a ranking, and no firm is recommended.

Outsourced software testing: a managed QA team on a retainer

Outsourced software testing hands test planning, execution and reporting to an outside team, by project or as a dedicated team. It fits a product team with a staging environment and more regression surface than its developers can click through. With one developer and no staging, the retainer is spent re-learning the app.

When you outsource software testing, the contract sets the model: QA Mentor, for one, lists seven, among them a dedicated team, fixed price and time and material. It earns its fee when releases come on a schedule, because each one brings the same regression pack back around and the testers already know the app.

It does not fit when there is one developer or none, no staging, and the app changes shape every week. Before you talk to outsourced software testing companies, get to a stable staging build, or the first weeks go on learning an app that keeps moving. If you do sign, insist on three things: test cases and bug reports in your own tracker, testers working on staging with anonymized data, and no production credentials in their hands.

Automated testing services: an agency builds the suite, then someone has to keep it green

Automated testing services pay an agency to write end-to-end and API tests and wire them into CI. The suite is only yours if it passes one test: it runs from your repository, in your CI, in an open-source framework, with no vendor login. Otherwise you bought a subscription, which is fine when that was the plan.

The work itself is plain. Engineers choose a framework, write end-to-end and API tests for your main journeys, wire them into CI, and then either hand the suite over or maintain it under a separate arrangement. Ask any seller of automation testing services to walk through these four parts before you sign; I call them the ownership test:

  1. The suite lives in your repository, not in the vendor’s account.
  2. It is written in an open-source framework. Playwright is a common choice, and describes itself as “an end-to-end test framework for modern web apps”.
  3. It runs in your own CI (GitHub Actions, for example) using secrets you hold: the test account, the staging URL, any API keys.
  4. A fresh clone runs it green against your staging app with every vendor login removed, and breaking one journey on purpose, say by making the checkout step return an error on staging, turns it red.

A correct handover passes all four. A suite that only runs on the vendor’s platform fails part 1 or 3, and a hollow suite, one that never touches the real code path, fails part 4. An automated software testing company whose work passes has handed you an asset you can keep running.

Automated testing companies fit when manual test cases already exist and releases are frequent. They do not fit when nobody on your side will own a red build, because an automated testing company can write the test but cannot decide whether the failure blocks your release. How the smallest useful suite is written is in how to write end-to-end smoke tests, and the kinds of tool an agency will propose are a separate subject: web application testing tools.

A QA-as-a-service platform or a managed tester community

These fit when you want coverage without hiring, your journeys are stable, and device or browser spread matters. They do not fit a one-off budget: QA Wolf’s platform bills by usage and Testlio’s by a platform fee plus a consumption fund, both for as long as you use them.

Before you sign, ask what you keep if you cancel: the test definitions, the run history, the bug reports. Read the answer in the terms, not the sales call. QA Wolf’s pricing page states, “Your tests are standard open-source Playwright and you can export them at any time.” That is the vendor’s claim, and it is the kind of answer to look for.

A freelance QA engineer, or a first QA hire

A freelancer fits a bounded job: write the first end-to-end tests, or test one release on real devices. A freelancer does not fit as the only safety net, because the tests need an owner after the contract ends.

A hire is priced by the salary row in the cost table, with benefits and tooling on top. Whether a first version needs a dedicated tester at all is its own question: whether a first version needs a QA engineer.

Nobody yet: what to do yourself this week

The honest option before buying any service is free. Write down the journeys that cost you money when they break, and click through each one before every release. What to check first, in order, is in how to test a vibe-coded app. If an AI tool also wrote your tests, read why tests pass but the app is still broken before trusting the green checkmarks.

A testing service does not replace reading the code. That is a different purchase, a code audit service, and a security test is another, covered in web application security testing services. Load testing and penetration testing are their own jobs, and so is beta testing with real users.

Questions to ask before you pay

Each question has one line on what a good answer sounds like. The good answers are my reading, and they all point the same way: at something you keep after the last invoice.

  1. 01 What exactly do I get back, and in whose system does it live? Good answer: test cases, bug reports and code in your tracker and your repository.
  2. 02 If I cancel, can I run the suite tomorrow from my own repository and CI with no login of yours? Good answer: yes, and here is how.
  3. 03 Which framework and language do you write tests in, and is it open source? Good answer: a named open-source framework in a language your developer reads.
  4. 04 Who maintains the tests when the interface changes, and at what monthly cost? Good answer: a named person or role and a written figure.
  5. 05 What is the minimum engagement, and what is the notice period? Good answer: both in writing, before any work starts.
  6. 06 What access do your testers need, and can all of it be staging with anonymized data? Good answer: staging only, and never production credentials.
  7. 07 How do you prove a test catches a bug: will you break a journey on purpose and show me the failing run? Good answer: yes, on a call or in a recorded CI run.
  8. 08 Have you tested an app an AI builder generated, and what did you find first? Good answer: a specific finding, not a slogan.

Question 7 comes from my audits, where a point-of-sale app had an 818-line test suite, green on every run, that never called the real sale-creation code; a green suite that never called the sale code tells it in full. The lesson I take from it is that a suite nobody has seen fail proves nothing yet.

A seller who cannot answer questions 2 and 7 in one sentence each is selling hours, in my view.

In the Production Hardening Sprint, deliverable 10.5 is verified this way: run the suite in CI and demonstrate that an intentional regression fails it.

Where the sprint fits

Deliverable 10.5, critical-path smoke tests, automates smoke tests for signup, login, the core product action, and payment flows, and its verify line is the one given after the questions above. Deliverable 10.7, auth and billing unit tests, adds unit tests around authorization and payment logic, including negative and edge cases; it is verified by running the tests and showing rejection of unauthorized access and incorrect billing transitions. Under 7.3, pull-request CI, we run linting, type checks, builds, and tests on every pull request. Under 7.2, protected production branch, we protect the main branch with required checks and controlled merge permissions. Deliverable 10.8 is AI repository guardrails. The sprint covers one codebase. Post-handover support is 14 calendar days of fixes for defects in the delivered sprint work, and 30 calendar days of async access for questions about the handover and architecture. Hosting, paid tools and API usage remain in your accounts. Each deliverable is listed in the published scope. Who keeps the suite green after that, your own team or one of the sellers above, is still your call.

Common questions about hiring a testing company

Will QA testers be replaced by AI?

Not on ONET OnLine’s current numbers. ONET rates projected growth for software quality assurance analysts and testers from 2024 to 2034 as “Much faster than average (7% or higher)”. AI tools can now write test scripts, and when the same tool writes both the code and the checks, a green run can prove little. I make no forecast of my own beyond that.

Is QA worth it in 2026?

For a small app, yes, starting with the first tests on signup, login and payment, which in my view are worth more than any test written after them. A QA retainer is worth it when releases are frequent and someone on your side owns the results; before that, a few end-to-end tests you own on those three journeys are the better first spend. Whether a first version needs a dedicated QA person is covered in the article on MVP development teams.

What are the top companies in software testing?

The ranked lists that answer this question are not neutral. On 2026-09-27, QASource’s USA list put QASource first, DeviQA’s rating put DeviQA first, and QualityLogic’s list put QualityLogic first; each list is published by the firm at its top. Use the eight questions and the four-part ownership test above to compare the firms you are actually talking to.

What are the four types of software testing?

ISTQB’s Foundation Level syllabus (v4.0.1, dated 2024-09-15) addresses four test types: functional, non-functional, black-box and white-box testing. In its words:

  • Functional testing “evaluates the functions that a component or system should perform”.
  • Non-functional testing “evaluates attributes other than functional characteristics of a component or system”.
  • Black-box testing “is specification-based and derives tests from documentation not related to the internal structure of the test object”.
  • White-box testing “is structure-based and derives tests from the system’s implementation or internal structure”.

The same syllabus describes five test levels, from component testing up, as a separate list; they are not the answer to this question. The source is ISTQB’s Foundation Level syllabus.

Is automation testing dead?

No. A scripted check that fails when a journey breaks is still how a team knows, on every release, that the journey works. What changed is who writes the scripts, and a script an AI tool wrote still has to be seen failing once before its green run means anything.