Skip to content
, home

How we measure results

How you’ll know it worked

Training is easy to sell and hard to prove. So every TestHubble option starts with a baseline and ends with a report you can take to your director.

We don’t promise “X% faster”. We promise what we’ll measure, and we report what we find, including what didn’t move.

your-team/measured
skillsAI adoptionflaky ratetime to fix

Baseline, check-in, results, follow-up

  1. before we start

    before we start

    Step 1: Baseline

    Each engineer takes a practical skills test in your framework, with your approved AI tools. We also pull the history you already have in your repo, CI, and tracker, ideally from the weeks before the program.

  2. halfway

    halfway

    Step 2: Mid-point check

    A short call with you: attendance, progress, early adoption, and anything we should adjust.

  3. at the end

    at the end

    Step 3: Results report

    The same kind of skills test with a different task of equal difficulty, the adoption numbers, and the team metrics compared with the baseline. Presented to you and, if you like, to engineering leadership.

  4. 30 days later

    30 days later

    Step 4: Follow-up

    We check the trend again after 30 days (and 60 for the 8-week options). Is it sticking, or fading?

Four layers, from the individual to the business

01 · the individual

Skills

Can each engineer design good tests, write automation that follows your conventions, use AI effectively, catch AI mistakes, and fix broken tests? Measured with a hands-on test before and after, plus a short self-assessment.

Same framework, same tools. A different task of equal difficulty at the end.

02 · the team’s habits

Adoption

Is the team actually using AI for testing work week to week? We track this with a two-question weekly pulse and an ai-assisted label on test PRs.

Example PR label:ai-assisted

03 · the work

Team outcomes

Using only data you already have. Volume is always read together with quality.

  • tests added per sprint
  • time to automate a typical case
  • automation backlog
  • flaky test rate
  • time to fix broken tests
  • review rework on test PRs

04 · the business

Business view

A conservative time-saved estimate built from your own numbers and cost figures, with every assumption shown, plus the benefits that aren’t time: less flakiness, a smaller backlog, and a playbook the team keeps.

hours saved × your cost per hourwith every assumption listed

results-report.pdfsample

Team outcomes: baseline and week 4

Example results report layout with placeholder values
Team outcomeBaselineWeek 4
Flaky test ratex.xx.x
Median time to fix broken testx.xx.x
Test PRs labelled ai-assistedx.xx.x

Context: what changed on the team during the program, and why a metric did or didn’t move.

Example layout. Your report uses your data.

What the report says, and what it doesn’t

A program like this isn’t a controlled experiment. Teams change, releases happen, and priorities shift. Your report lists what changed during and after the program next to that context, so you can defend it rather than oversell it. If a metric didn’t move, the report says so and suggests why.

This is development, not surveillance

Engineers do better work when they know how their data is used. These rules are agreed with you at kickoff and explained to every participant before the first test.

Each engineer

  • Gets a written notice at kickoff: what we measure, why, who sees what, and how long we keep it.
  • Sees all of their own results: scores, instructor notes, and a growth plan.
  • Can review and correct their results before the final report goes to the manager.

The manager

  • Sees team totals for every measure.
  • Sees, per person: attendance, lab completion, before/after level on each skill, and capstone status.
  • Does not see detailed instructor notes unless the engineer agrees.

For everyone

No rankings.
We never rank engineers against each other.
Not a performance review.
Our agreement asks that program results are not used as the only basis for HR decisions.
Short retention.
Skills-test recordings and AI chat logs are used only for scoring and deleted within 30 days of the final report. Reports belong to you; we keep only anonymized totals.

Want to see what we’d measure for your team?

On a call we’ll look at which metrics you already have and which ones we’d set up at the start.

you already haveset up at the start