Trial & Experiment Log
Run process trials that prove something — a hypothesis written before you start, a measured baseline, one variable changed, a success threshold agreed in advance, and an honest conclusion with the sample size printed next to it. Nothing is uploaded.
Version 1.0.0 · Updated Aug 8, 2026
Overview
Frequently asked questions
How does the Trial & Experiment Log licence work?
It is a one-time purchase for a downloadable tool — no subscription. You buy it once and the file is yours to keep and use.
Can I try the Trial & Experiment Log before buying?
Yes. Use the Try online button for a fully interactive demo with sample data already loaded — nothing to install and nothing is saved.
Can I import my data from a spreadsheet?
Yes. Use the Spreadsheet template button to save a CSV with the right headings, fill it in Excel or any spreadsheet, then Import spreadsheet to load it back. The file is read in your browser — nothing is uploaded.
Does my data stay private?
Yes. The tool is a single HTML file that runs entirely on your computer and makes no network requests, so nothing you enter is ever uploaded or shared.
Do I need Excel or any other software?
No. It replaces the spreadsheet template entirely: open the file in your browser (Chrome, Edge, Firefox or Safari) on Windows, Mac, Linux or a tablet, and start working.
How to use Trial & Experiment Log
The complete in-tool guidance, reproduced here so you can read it before you download.
What this tool does
CM8-313 is a register of process trials. One row is one trial: the hypothesis written before you started, the baseline you measured, the one thing you changed, the target agreed in advance, the result, the observations behind each side of it, what else was going on, and what you decided.
It has three neighbours and replaces none. The Kaizen Tracker is the pipeline of improvement ideas; a trial is what you run when an idea is uncertain enough that acting on faith would be reckless. The Lessons Learned Register captures what finished work taught you, in words. The 8D Problem Solving Record drives a defect to a root cause and a permanent fix. This one asks only: did the change work? It all runs inside this single file — no account, no upload, no network request.
Why most workplace trials prove nothing
The typical trial goes like this. Somebody suggests a change. It is tried for a few days, usually by the person who suggested it, on the easy jobs, while they are paying more attention than anybody will pay afterwards. Nobody measured what the old way produced, so the comparison is against a remembered number. Nobody wrote down what "better" would look like, so the finish line moves to wherever the result landed. Something else changed in the same fortnight and nobody noted it. Six observations are set against a vague impression of the last six months, one person announces it seemed better, and because that person wanted it to be better, it is adopted. Two months later output is exactly where it was, and the change disappears the week its champion goes on leave. Nothing was learned, because nothing was tested — and the organisation now believes one more thing for no reason.
The remedy is not statistics. It is four small things, written down before the trial starts, where they cannot be edited afterwards to fit the answer.
The four things a trial needs
- A hypothesis, stated first. If we change X, then Y will improve, because Z. The "because" is the part people skip and the part that earns its keep — it names the mechanism, and a wrong mechanism teaches you something even when nothing moves.
- A measured baseline. Not a memory, not a standard time. Measure over long enough to include an ordinary bad day, and record the period and the observations it covered.
- One variable changed. Change the feed rate and the insert grade in the same week and you learn that the pair did something, without learning which. The "what is being changed" field is one line on purpose: if it will not fit, split the trial in two.
- A success threshold agreed in advance. The target makes the trial falsifiable. Deciding afterwards what counts as success is how honest people fool themselves, so the tool will not let you conclude adopt or reject without one.
Sample size, in plain language
Every measure wanders — with the operator, the material, the time of day. Take six readings from the old way and six from the new, and the difference between those handfuls is mostly wander. Take more and it averages out, so more of what is left is real: the fewer observations you have, the less the difference means, however big it looks.
The minimum sample size setting exists to stop a two-part trial being reported as a result. When either side falls below it, or has none recorded, the evidence column reads Underpowered and the trial joins the credibility table — which lists every reason somebody could refuse to believe a result here. Twenty a side suits a measure that varies moderately; raise it for anything noisy. A blank counts as not yet demonstrated: an unrecorded sample size is not evidence of a large one.
This tool does not compute statistical significance. No p-value, no confidence interval, no power calculation — it compares two numbers and counts the observations behind each. A decision carrying real money, safety or regulatory weight deserves a properly designed test, analysed by somebody competent to do it.
Confounders
A confounder is anything else that changed during the trial and could explain the result on its own: a new crew, a different part number, a machine service, or simply that people work differently when watched. They are not a sign of a bad trial — a working factory will not hold still for you.
Writing them down costs a sentence and is what makes a trial honest: it stops you claiming more than the trial supports, and it suggests how to run the repeat. When an adoption carries recorded confounders the evidence reads Confounded — not because the decision was wrong, but because the attribution is uncertain and the register should say so. The flag targets adoptions, because that is where an unexamined confounder does damage. Anything in the box counts as one; if you found none, leave it empty and say so in the notes.
Noise and the meaningful-change threshold
Set the threshold from how much your measure moves between two ordinary periods with nothing changed. If scrap wanders between 3.6% and 4.4% week to week for no reason anybody can name, a 5% movement in a trial is indistinguishable from noise. Movements below it are grey on the change chart; at or above, green; negative, red. Wanting to lower it so a favourite trial passes is the threshold working.
Rejecting trials openly
A register where every trial was adopted is not a register of trials. It is a record of decisions already taken, dressed up. Real trials fail often — the mechanism was not the binding constraint, the saving was smaller than the disruption, the change fixed one station and broke the next. Recording those honestly is what makes the adoptions worth anything: a reader who can see you sometimes say no has a reason to believe you when you say yes. So reject openly, keep the row, and note what you now think was happening. If the rejected and inconclusive tile reads zero on a register of any size the tool says so: a log with no rejections is a log nobody is being honest in. Inconclusive is legitimate too — it names the confounder the repeat must design around.
Adoption — where the change lives
A trial that works and never reaches a standard has changed nothing. The bench stays where the trial left it, the setting holds until the next service, the sequence survives until the person who ran the trial takes a fortnight off — then everything returns to the documented method, which is what people fall back on.
The adopted into column is the fix, and it is enforced: you cannot conclude a trial as adopted without naming the standard, work instruction, programme revision or machine setting the change now lives in. The tile counting adopted trials with nothing named is the gap where improvements evaporate, and the sample register carries one: a packing layout everybody agrees works, still absent from the work instruction.
One trial at a time
Four trials at once in the same area is one trial with four variables: nothing can be attributed and the fortnight produces no usable knowledge. Run them in sequence — one variable, long enough to gather the sample, a conclusion recorded, then the next. The timeline makes overlap visible, and catches the trial nobody ended: a half-filled bar left of the dashed today marker has its data collected and nobody willing to look.
The formulas
Change % (higher is better) = (trial value − baseline value) ÷ |baseline value| × 100 Change % (lower is better) = (baseline value − trial value) ÷ |baseline value| × 100 Target met = trial value at or beyond the target, in the improving direction Evidence = Underpowered if either sample size is missing or below the minimum; otherwise Confounded if an adoption has confounders recorded; otherwise Reasonable
Change percent is direction-aware, so a positive number always means improvement whichever way the measure travels. It shows "—" when the baseline or trial value is missing, and when the baseline is zero, because a percentage change from zero is undefined. A blank is never read as a zero.
The spreadsheet workflow
- Spreadsheet template saves a CSV whose headings are exactly this tool's column labels, with a guidance row showing what each expects — the date format, and the accepted direction and conclusion values.
- Fill in the design columns first and add results as trials finish. Delete the guidance row before saving, and keep the CSV format.
- Import spreadsheet reads it back, matching columns by heading, so order does not matter and extras are ignored. Rows failing a check — an adoption with no target, an adoption with nowhere to live, an end date before the start — are skipped and reported by row number.
Nothing is uploaded, and importing adds to what is here rather than replacing it.
FAQ
How long should a trial run? Long enough to collect the minimum sample under normal conditions, and to cross whatever cycles the process has — a shift rotation, a material batch.
The measure improved but the target was missed. Adopt or reject? Beating the noise threshold while missing an ambitious target is often worth adopting; missing a carefully set one is usually worth extending. Write the reasoning in the notes.
Should an adopted trial be re-measured? Yes, three months on — the cheapest audit there is. If the improvement has gone, either the change never reached the standard, or the standard is not followed.
Saving your work
Trials, settings and the report header are saved to this browser's local storage as you type — one browser on one computer. Treat Export .json as the real save, which Import .json restores anywhere; Export CSV gives you the register for spreadsheet work; Reset asks twice, then erases everything stored here.
Accuracy & disclaimer
The arithmetic is simple and the tool does it faithfully. What it cannot do is larger: it cannot tell whether the baseline period was representative, whether the two populations were comparable, whether the measure was taken the same way both times, or whether the thing you did not write in the confounders box is what moved the number. It does not test statistical significance. A small sample with a big number in it is still a small sample. This is a record-keeping and discipline aid for process trials — not a statistics package, not a validation protocol, and no substitute for a properly designed experiment.
Related tools
5S Audit
Run a 5S workplace audit — score Sort, Set in order, Shine, Standardise and Sustain checkpoint by checkpoint from 0 to 4, track the audit score over time, and turn every low score into a corrective action with an owner and a date. Nothing is uploaded.
APQP Planner
Plan a product launch across the five quality-planning phases — one row per deliverable, with owners, planned and actual dates, gate deliverables, risk notes, phase completion and a Gantt of the whole programme. Nothing is uploaded.
Run 5 Why root cause analyses: state the problem, walk the why chain, name the root cause, then track the countermeasure through to verified. Runs entirely in your browser — nothing is uploaded.
Run 8D problem-solving reports discipline by discipline — team, containment, verified root cause, corrective action, prevention — with a board that shows exactly where each report is stuck. Nothing is uploaded.