Services
CRO & A/B Testing
Test design, sample size, and statistical significance done properly — including honesty about how often a test finds nothing at all.
- of tests
- 74.2%
- Clients served
- 60+
- Ad spend managed
- $20M+
- Years operating
- 3
found no statistically significant difference at all. Only 17.4% found a significant winner. That's the honest base rate for A/B testing, and any agency promising a win on every test is setting an expectation the data doesn't support.
2,408-test study, Visionary Marketing, 2026
Across a 2026 study of 2,408 A/B tests, 8.4% reached statistical significance with the variant performing *worse* than the control (Visionary Marketing, 2026 dataset). A test that reaches significance is not the same as a test that won, and the difference is money.
Most CRO programs fail at the setup stage, long before a test ever launches.
ThinkMedia runs CRO and A/B testing for companies spending $5,000 or more a month on marketing. We design tests properly, report the ones that fail exactly as clearly as the ones that win, and won't call a result before the sample size supports it.
Four mistakes show up constantly:
01
Peeking. A test shows a promising result on day three, someone stops it and calls a winner, and the result is statistically meaningless — it was never run to the sample size the test needed to reach a trustworthy conclusion.
02
No pre-defined sample size or minimum detectable effect. Without deciding in advance how big an effect matters and how many visitors are needed to detect it, any “significant” result midway through a test is just noise that happened to look like a pattern.
03
Testing too many things in one variant. Headline, image, and button color all change at once, so a win or a loss can't be traced to a specific cause, and the next test starts from zero learning instead of building on the last one.
04
Declaring victory after a partial business cycle. A test run for three days misses weekday-versus-weekend behavior entirely, and a real result needs at least one full business cycle, typically seven to fourteen days, to hold up.
How this runs
What you get.
- Reporting
- a written result for every test, including the ones that didn't win, with the effect size and confidence level stated plainly rather than summarized as a vague “improvement.”
- Access
- your testing tool and results data live under your own account from day one. On exit, you keep every test result and the full historical log.
- Communication
- a named strategist who designs and reads your tests directly, reachable with a one-business-day response commitment.
- Terms
- thirty days' notice, no annual minimum, flat fee — never billed per “win,” since that would create an incentive to call tests early.
What's included
The actual thing we do, and how often.
monthly planning cycle
Test prioritization and hypothesis design
each test starts with a written hypothesis and a stated reason to believe it will move the metric, not a guess dressed up as a test.
per test
Sample size and Minimum Detectable Effect calculation
calculated before launch, so everyone agrees in advance how long the test needs to run and what size of effect would actually be worth shipping.
per test
Test implementation and QA
variants built and checked in a staging environment before launch, including a validation pass that tracking fires correctly for both control and variant.
ongoing per test
Statistical monitoring without peeking
results tracked against the pre-defined sample size, with no early calls and no stopping the moment a result looks promising.
at test conclusion
Test result reporting
a written report stating whether the result was significant, what the effect size was, and what we're testing next because of it — including tests that found nothing, reported with the same clarity as the ones that won.
How we work
Four phases, always in this order.
Funnel Teardown
Days 1–10
We map your funnel and identify where the biggest, most testable drop-offs actually are, using existing analytics data rather than guessing which page deserves attention first.
Test Blueprint
Days 7–17
A prioritized test roadmap, each test with a written hypothesis, target metric, and calculated sample size before anything gets built. This is where we agree on what a real win would look like, before we're looking at results that could bias that judgment.
Build
Ongoing, per test
Each test is built, QA'd, and launched on its own schedule. What we deliberately do not do: call a test early because the dashboard looks good on day four. A test ends when it hits its pre-calculated sample size, not when the number we want to see shows up.
Compounding
Ongoing
A continuous testing calendar, each test informed by what the last one taught, whether it won, lost, or found nothing — since even a null result usually rules something out and narrows what's worth testing next.
Honest scoping
Who this is for, and who it isn't.
You're a good fit if you have enough traffic to reach statistical significance within a reasonable timeframe — generally a few thousand visitors per variant per month at typical conversion rates — and an existing funnel worth optimizing rather than building from scratch.
You're not a good fit yet, and we'll say so before taking the engagement, if your traffic volume is too low to reach significance within a normal testing cycle, since a test that would take a year to reach sample size isn't a good use of your budget yet. If your conversion funnel doesn't exist yet, or converts so rarely that only a handful of conversions happen a month, build the funnel first — testing optimizes something that's already working, it doesn't create it.
| Opinion-based redesign | A/B testing | |
|---|---|---|
| Decision basis | Taste, HiPPO | Measured visitor behavior |
| Risk if wrong | Full traffic affected immediately | Limited to the test split |
| Confidence in result | None | Statistically calculated |
| Learning captured | Rarely documented | Every test logged, win or lose |
| Typical outcome rate | Unknown | ~17% clear win, ~74% no difference |
Asked and answered
Common Questions
ThinkMedia charges a flat monthly fee for ongoing testing programs, not a per-test or per-win fee, starting for companies spending $5,000 or more a month on marketing overall. Pricing is scoped after the funnel teardown.
It depends on your traffic and baseline conversion rate, calculated per test before launch — typically two to six weeks to reach a pre-defined sample size, run across at least one full business cycle to account for weekday and weekend behavior.
Not every test wins, and any agency promising that isn't being honest with you. Across a large 2026 study, only about 17% of tests found a significant winning variant, while roughly 74% found no significant difference at all — that's the normal, expected outcome distribution for rigorous testing.
You own your test results and data at all times. Testing tools are set up under your own account from day one, and the full historical test log stays with you if you leave.
There's no fixed number, since it depends on your baseline conversion rate and the effect size worth detecting, but generally a few thousand visitors per variant per month is the practical floor for reaching significance within a normal testing cycle.
It gets reported exactly as clearly as a win would be, with the effect size and confidence level stated plainly, and it usually narrows what's worth testing next — a null result rules something out, which is real information, not a wasted test.
A named strategist assigned at kickoff, the same person who calculates sample sizes and reads results, reachable directly rather than through a rotating analytics team.
ThinkMedia has no minimum contract length. Thirty days' notice ends the engagement in either direction, with no annual minimum.
Proof
Results.

3.9x
Blended ROAS
from 1.6x
3.9x ROAS and $265k in Monthly Revenue for a Los Angeles Apparel Brand
We took Vespera Apparel from a single undifferentiated Meta campaign and a 14% repeat purchase rate to a segmented account and real retention infrastructure — nearly tripling blended ROAS along the way.

3.8x
Blended ROAS
from 1.7x
214% Revenue Growth and 37% Lower Cost Per Purchase for an Austin Skincare Brand
We took Velora Beauty from a hero-product-only account with fatigued creative to a routine-building brand — and nearly doubled repeat purchase rate in the process.

3.6x
Blended ROAS
from 1.9x
3.6x ROAS and CAD $325k Monthly Revenue for Havenwood Home
We took Havenwood from product pages that couldn't answer a considered buyer's real questions to a full storytelling system for solid-wood furniture — and more than doubled repeat purchase rate along the way.
Get in touch
Tell us what you’re running.
You send the details
Channels, monthly spend, and the part that is not working.
We look at the accounts
Sixty to ninety minutes inside them, before we say anything.
You get the findings
A 45-minute call covering everything — including what you can fix yourself.
Keep going
The rest of it.
- ConversionLanding Page DesignBuilt specifically for paid traffic: fast, message-matched, mobile-first, and tracking baked in before the first visitor arrives.Read more
- ConversionTracking & AttributionGA4, server-side tagging, Conversions API, enhanced conversions, and Consent Mode — built so your ad platforms optimize against what actually happened.Read more
- ConversionSEO ServicesRankings still matter. They're just no longer the whole job — brands cited inside an AI Overview earn 35% more organic clicks than brands ranked but not cited on the same query.Read more
- Performance mediaGoogle Ads ManagementFor companies spending $5,000+/month who need to know which campaign actually made the revenue, not which one the platform says did.Read more
- Performance mediaMeta Ads ManagementFor companies spending $5,000+/month who need creative volume and clean signal, not another campaign type change to react to.Read more
- Performance mediaLinkedIn AdsFor B2B companies who need to reach a buying committee by job title, seniority, and company — not by guessing from someone's interests.Read more