Cookie preferences

Essential cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.

Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.

Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.

Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.

Skip to main content
One year, four agents: How AhTestHub evolved beyond automation
Person using a laptop for QA testing

Main Points

  • AI is producing code faster than developers can trust it, and QA is the discipline that closes that gap.
  • AhTestHub has evolved into four AI agents covering test generation, automation, accessibility and performance.
  • Every agent's output is checked by a human. No unreviewed AI output reaches production.
  • Engineers now focus on risk, edge cases and intent rather than boilerplate, prioritising quality over quantity.
  • The real question isn't whether AI can write software but who's checking it, and at All Human, humans are.

Lately there’s been a lot of conversation around AI and how best to regulate it without inhibiting progress. That is great and needed, and the more this discourse moves into and consumes the public realm, the better, since these conversations typically happen behind closed doors. However, as someone who has worked in QA for many years, sometimes I think these current conversations overshadow the fact that QA has been a constant force and a fundamental part of technology for a long time. AI did not introduce the concept of quality checks and guardrails; they’ve been there all along.

Here at All human, the QA process is rigorous, well established, and strictly adhered to so we can ensure any product or service delivered to our clients performs at the highest level. There are five steps: groom, test the design, execute, regress, and finally sign-off. Five steps that we conduct for every release, no shortcuts, no exceptions. 

And it has worked brilliantly. 

But with the ever-increasing use of AI to generate code, it was time to rethink our testing process. 

Why?

Primarily because the code is arriving faster than anyone's confidence in it. AI now accounts for 42% of all committed code, yet 96% of developers say they don't fully trust that what it produces is functionally correct — and fewer than half of them verify it before it's committed. Stack Overflow's 2025 survey points the same way: developer trust in the accuracy of AI output has fallen, not risen, as adoption has climbed.

That's a huge gap. And the thing that closes that gap — verification, evidence, governance — happens to be exactly what QA does.

So the question we asked was not " does AI replace our QA process?" It was "our QA process is the thing everyone suddenly needs more of; how do we scale without diluting it?"

And we have.

The AhTestHub

What started as an automation framework has evolved over the past year into something closer to a set of specialised AI agents working in sequence. Each has one clear job which a human engineer checks before anything moves forward.

And for clarification, when I say agent I am referring to an AI worker that plans and produces work within a narrow, well-defined scope, and then hands that work to a person. It does not mean software that decides on its own what "good" looks like.

There are four jobs and four checkpoints:

Test generation - the agent drafts, the engineer decides.

We feed in project requirements and the agent synthesises test scenarios across viewports. The QA engineer then edits, standardises, and approves or rejects every generated case before it becomes a story-linked test case ready for execution. What used to start as a blank page now starts as an edit.

Automation build — from approved case to running coverage in hours, not days

Once a case is approved, it's compiled into a structured, executable test — Playwright for web, Appium for native mobile — following our existing conventions, with self-healing locators to reduce flakiness from UI churn and accessibility assertions injected by default. It's wired into continuous integration (CI) from day one.

Accessibility — no longer a separate pass at the end

We've long advocated for digital accessibility and have worked on multiple award-winning accessible sites, so this is the part I'm proudest of. WCAG 2.2 AA validation is baked into every generated suite by default, categorised by severity across contrast, ARIA roles, keyboard order, heading levels and alt text.

Performance — continuous benchmarking, not a quarterly audit

Lighthouse runs on every URL, every release, trending release over release rather than captured once and filed. On a recent build that meant scores of 98 performance, 100 accessibility, 96 best practices and 100 SEO — and, more usefully, a flag the moment any of them starts to slide.

And running through all four is one fundamental,non-negotiable rule: no unreviewed AI output reaches production. AI generates. Humans approve. That principle is not a caveat we added at the end, it's the reason the rest of it is safe to use at all.

AhTestHub - Quality sits at the centre.
AhTestHub - Quality sits at the centre.
What QA automation means for the team

The honest answer is that the work got more interesting.

Our engineers spend less time writing boilerplate and chasing broken locators, and more time on the things that need human, expertise:

  • what's genuinely risky in this release, 
  • what an edge case actually looks like for this user, 
  • whether a generated scenario has understood the intent behind the acceptance criteria or just its wording. 

We also stopped optimising for raw test coverage and started optimising for risk coverage. Conducting more tests is not the goal, fewer, better-targeted tests against the areas most likely to break is.

Quality not quantity.

Built to keep evolving

The AhTestHub roadmap is less a fixed destination and more about preparing for what’s next.  

Currently we are looking at: 

  1. A technology-agnostic foundation — no commitment to a single stack.
  2. Continuous scanning — monitoring emerging tools and approaches as they mature.
  3. Adaptation by project need — choosing what best supports each individual engagement.
  4. Flexible standards — staying rigorous without becoming rigid.
The bottom line

We believe AI has real applications across business transformation, and QA is one of the clearest. And because we've always prioritised quality, it made sense that we'd use AI to sharpen it rather than shortcut it.

Collectively we’ve spent this year asking whether AI can be trusted to write software. The more useful question is who's checking. We're making sure the answer at All human is: we are, and now at a scale that matches the pace.

Social

Enjoyed this article? Share it with someone else.

You might also like