Most teams guess at why documents sit unsigned. A/B testing replaces the guess with evidence: you send two versions of a signing request, change exactly one thing between them, and measure which gets more signatures. Done properly, it turns "I think the subject line is the problem" into a number you can act on.
Done carelessly, it produces confident conclusions from noise — which is worse than not testing at all. This guide covers both sides: how to run a valid test, and how to know when your result means nothing.
Key takeaways
Change one variable per test. Two changes means you learn nothing about either.
Sample size decides what you can detect. With 100 recipients per variant you can only spot very large differences.
Run for at least 3–5 business days. Signature completion has a long tail.
Subject line, sender name, and reminder timing are the highest-leverage starting points.
Don't A/B test legal wording, disclosure text, or anything that changes what the signer is agreeing to.
Transactional signing emails have deliverability constraints that marketing emails don't.
What you need before you start
A baseline. Pull your current completion rate, open rate, and median time-to-sign from the last 90 days. Without this you can't tell improvement from normal variation — completion rates swing week to week on their own.
Enough volume. This is the constraint most teams underestimate. See the sample size table below before planning anything.
One clear primary metric. Completion rate is usually right, because it's the outcome you actually care about. Open rate is a means to it, not a goal — a subject line that boosts opens but not signatures has told you something, but not something useful.
A way to split recipients randomly. Random assignment is what makes the comparison valid. Splitting by "first half of the list, second half" silently sorts by signup date, region, or account size.
A written hypothesis. "Naming the deadline in the subject line will increase completion, because recipients currently deprioritise requests with no visible urgency." Writing this down before you send stops you from reinterpreting the result afterwards to fit whatever happened.
How many recipients do you need?
This is the part most guides get wrong. Sample size depends on how big a difference you're trying to detect — smaller effects need much larger samples.
Rough figures for a 50% baseline completion rate, 95% confidence, 80% power:
Lift you want to detect | Recipients per variant |
|---|---|
50% → 55% (5 points) | ~1,600 |
50% → 60% (10 points) | ~400 |
50% → 65% (15 points) | ~165 |
50% → 70% (20 points) | ~95 |
Read this as a warning, not a target. If you send 200 documents a month, you can only reliably detect changes of 15 points or more. Anything subtler is invisible to you, and a small apparent difference in your dashboard is almost certainly noise.
If your volume is low, you have three honest options: accumulate results across several months of sends, only test changes big enough to produce large effects (rewriting the entire email, not tweaking a word), or skip testing and apply known-good practices from signature field placement strategies and forms and fields that get signed instead.
Step 1: Pick one variable
Ranked roughly by impact per unit of effort:
Subject line — affects whether the request is opened at all.
Sender name — a named person usually outperforms a generic system address. If you're sending from a no-reply address on a shared domain, fixing that comes before testing it. Sending from your own business domain affects both trust and inbox placement.
Reminder timing — often the single largest lever, and underused.
Signing order — sequential versus parallel changes total turnaround substantially on multi-signer documents. Common signing workflow patterns will tell you which ones apply to your documents.
Document length and field count — fewer required fields, faster completion.
Channel — email only versus email plus SMS. See omnichannel signing explained.
Start with the subject line. It's fast to change and its effect shows within hours.
What not to test: clause wording, disclosures, consent language, or anything that changes the legal substance of what the signer agrees to. Also avoid manufactured urgency — a fake deadline may lift short-term completion while creating a genuine problem if the agreement is later disputed. Compliance constraints outrank optimisation. [LEGAL REVIEW]
Step 2: Build the two versions
Version A is your current request, unchanged. Version B changes one element.
Example:
A: "Please sign your contract"
B: "Your contract is ready to sign — takes about 2 minutes"
Everything else stays identical: same send time, same body copy, same document, same signer order. If B also has a different sender name, the test is ruined and you won't know it.
Send both at the same hour on the same day. Time of day has a large effect on open rates and will swamp whatever you're measuring if you let it vary.
For personalisation tests specifically, personalised signing experiences and completion rates covers which elements tend to move the needle.
Step 3: Check deliverability before you blame the copy
Signing requests are transactional email. They're expected to land in the inbox reliably, and they carry a different set of rules than marketing sends — see transactional versus marketing email for where the line sits.
Two practical consequences:
Authentication comes first. If your domain lacks proper SPF, DKIM, and DMARC records, a share of your requests never reach the inbox at all, and every test you run is measuring deliverability noise rather than copy. Why business emails land in spam covers the fix. Do this before testing anything.
Urgency language carries risk. "ACTION REQUIRED", "URGENT", deadline pressure and heavy capitalisation are the same patterns spam filters look for. A variant can win on opens among recipients who received it while quietly reducing delivery overall — and your open-rate denominator won't show that.
Step 4: Run it, and leave it alone
Run for a minimum of 3–5 business days; a full week if your recipients span multiple time zones. Signature completion has a long tail — a meaningful share of signers act on day 3 or later, and stopping on day 1 systematically favours whichever variant appealed to fast movers.
Decide your stopping point before you start and don't move it. "Peeking" at results daily and stopping the moment one variant pulls ahead is the most common way to generate a false winner: with enough looks, random variation will eventually produce an apparent lead.
Avoid holiday periods, quarter-end, and anything else that makes the week unrepresentative.
Step 5: Judge the result honestly
Run your two numbers through a two-proportion significance calculator. Three outcomes are possible, and two of them are not wins:
Significant difference → adopt the winner.
No significant difference → the change doesn't matter at your volume. This is a real, useful result. Record it and move on.
Difference that looks large but isn't significant → treat as no result. This is where most false wins are born.
Then check two secondary things. Did time-to-sign improve or worsen? And did support enquiries go up — a variant that confuses people can raise completion while creating work elsewhere.
Segment the result if you have volume for it. A variant that wins overall can lose badly with one customer type. Just be aware that slicing results after the fact multiplies your chances of finding a fluke, so treat segment findings as hypotheses for the next test rather than conclusions.
Worked example (illustrative figures)
A company sending software licence agreements has a baseline completion rate of 52% and sends roughly 1,000 agreements a month. They want to test whether naming a deadline helps.
At 500 recipients per variant, the table above says they can detect a difference of roughly 9 points or more. Anything smaller is out of reach, and they accept that before starting.
A: "Your licence agreement" → 52% completion
B: "Your licence agreement — please sign by Friday" → 61% completion
Nine points, 500 per arm: this clears significance. They adopt B, then run a second test on reminder timing (24 hours versus 72 hours after send), which is the natural follow-up once the opening message is settled.
Note what this example does not claim: a revenue figure. Attributing a specific dollar amount to a subject line requires assumptions about contract value, close rate, and what those signers would have done anyway. If you report ROI, show the assumptions.
Common mistakes
Two variables at once. You learn nothing about either.
Stopping early on a favourable reading. Set the end date up front.
Calling a small gap a win. A 3-point difference on 80 recipients is noise.
Testing copy while authentication is broken. Fix the domain records first.
Ignoring the long tail. Late signers change results more than people expect.
Never writing anything down. Without a log of hypothesis, result, and decision, the same test gets run twice a year apart.
What to do at low volume
If you send fewer than about 200 documents a month, formal A/B testing won't give you reliable answers. Better uses of your time:
Apply established practice on field placement and form design.
Reduce required fields wherever the information isn't strictly needed.
Start from a pre-built template rather than a fresh document each time — consistency alone removes a class of friction.
Fix reminder timing, which usually helps regardless of audience.
FAQ
What is A/B testing in signing campaigns?
Sending two versions of a signing request to randomly split halves of your recipients, changing one element, and comparing completion rates. The goal is to identify what genuinely drives signatures rather than relying on assumption.
How many recipients do I need?
It depends on the size of the difference you want to detect. Roughly 400 per variant for a 10-point lift from a 50% baseline; around 1,600 per variant for a 5-point lift. With under 100 per variant, only very large differences are detectable.
What should I test first?
The subject line, then sender name, then reminder timing. These are quick to change and affect the earliest steps of the funnel, where most drop-off happens.
How long should a test run?
At least 3–5 business days, or a full week for multi-time-zone audiences. Signature completion has a long tail, and ending early biases the result toward fast responders.
Can I A/B test legal or HR documents?
You can test how a document is delivered — subject line, sender, reminders, field layout. Don't test the legal content, disclosures, or consent language. Those need legal review, not optimisation. [LEGAL REVIEW]
Will testing subject lines hurt deliverability?
It can. Urgency words, capitals, and pressure phrasing trigger spam filters. Confirm your SPF, DKIM, and DMARC are correct before testing, and check delivery rate alongside open rate.
What if neither version wins?
That's a valid result: the element you tested doesn't matter at your volume. Log it and test something with larger expected impact.
How do I avoid bias?
Use genuine random assignment rather than splitting the list in order, send both variants simultaneously, keep everything else identical, and fix your end date before you start.
