Writing your own tests

Lesson 0009 · after 0008 · reading, then a 45-minute drill graded by planted bugs · ~55 minutes

Every code block, every compiler message, and every terminal session on this page was produced by running it today. Nothing is written from memory. Where a cargo test block is quoted, the Compiling / Finished lines and the empty Doc-tests section are cut and nothing else. The demo domain is a thermostat, defined in full two sections down — your project is a task CLI, so nothing here pastes in. Translating is the work.

Where 0008 left you, and the bug that 46 tests could not see

The library half of 0008 landed cleanly. All 46 tests pass, tally is a real generic function written from its signature, load is one collect::<Result<Vec<Task>, TaskError>>()?, remove_completed uses Vec::retain, and count_by_priority is a one-line delegate. Iterators and HashMap are produced, not recognised.

Three of the drill's checks still failed, and the interesting thing is where they failed. Here is your stats and clear, run today against your own crate:

$ run stats
high 	 1
low 	 1
medium 	 1
$ run clear
$ run list
2 [todo] call bank (medium)
3 [todo] water plants (low)

The priorities print in the wrong order — high, low, medium, because main.rs loops over [Priority::High, Low, Medium] — and clear prints nothing at all, throwing away the usize that remove_completed went to the trouble of returning. Both are real defects a user would notice in the first minute. Both sat behind 46 green tests.

That is not bad luck, and it is not because you were careless. It is structural: every one of those 46 tests lives in tests/, every one of them talks to the library, and run lives in src/main.rs, where no test in tests/ can reach it. The untested thing is the undone thing — this is the third lesson in a row where the one requirement no test could see is the one requirement that was not met. Today you close that loop from both ends: you learn to write tests, and you move the code that was unreachable into a place where a test can reach it.

The demo domain, in full

Every example on this page runs against one small crate, so it is worth reading the whole thing once before the examples start. It is a thermostat that refuses illegal targets. Three things to notice as you read: the field target is private, the helper capped has no pub, and new panics while set_from returns a Result — the page needs both to show you both ways of testing failure:

// /tmp/heating/src/lib.rs — created with `cargo new --lib heating`
pub const MIN: i32 = 5;
pub const MAX: i32 = 30;

#[derive(Debug, PartialEq)]
pub struct Thermostat {
    target: i32, // degrees celsius, always inside MIN..=MAX
}

impl Thermostat {
    // panics if the target is outside the legal range
    pub fn new(target: i32) -> Thermostat {
        if target < MIN {
            panic!("target must be at least {MIN}, got {target}");
        } else if target > MAX {
            panic!("target must be at most {MAX}, got {target}");
        }
        Thermostat { target }
    }

    pub fn target(&self) -> i32 {
        self.target
    }

    // never leaves the legal range, however big `by` is
    pub fn warmer(&mut self, by: i32) {
        self.target = capped(self.target + by);
    }

    pub fn is_heating(&self, room: i32) -> bool {
        room < self.target
    }

    // "21" -> Ok, "hot" or "99" -> Err
    pub fn set_from(&mut self, text: &str) -> Result<(), String> {
        let degrees: i32 = text
            .trim()
            .parse()
            .map_err(|_| format!("not a number: {text}"))?;
        if degrees < MIN || degrees > MAX {
            return Err(format!("out of range: {degrees}"));
        }
        self.target = degrees;
        Ok(())
    }
}

// private helper: no `pub`, so only this file can call it
fn capped(degrees: i32) -> i32 {
    degrees.clamp(MIN, MAX)
}

One naming convention, because it is the only way to read the examples without guessing: Thermostat (capitalised) is the type, t is always a value of it, and room is always the current room temperature rather than the target. The mapping onto your crate is loose on purpose — this is a different program, not a template. What transfers is the shape of a test, not its subject.

Part 1 — A test is a function that fails by panicking

The whole mechanism is one attribute. Put #[test] on a function that takes no arguments and returns nothing, and cargo test builds a second binary out of your crate, runs every such function, and reports on each one. The book's definition is worth reading slowly, because the second half of it is the part people never internalise:

Tests fail when something in the test function panics. Each test is run in a new thread, and when the main thread sees that a test thread has died, the test is marked as failed.

Book: 11.1 — How to write tests

So there is no assertion framework here and no special test runtime. Panicking is the failure protocol. Every assertion macro you are about to meet is a thin wrapper that panics when its condition does not hold, which is why .unwrap() in a test body is not a code smell the way it is in main — an unwrap that fires is a test that fails, with the message you wanted anyway.

Here are two tests against the thermostat, and the output they produce. Read the output as carefully as the code, because the failure format is the thing you will actually spend your time reading:

#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn a_new_thermostat_keeps_its_target() {
        let t = Thermostat::new(20);
        assert_eq!(t.target(), 20);
    }

    #[test]
    fn warmer_never_passes_the_maximum() {
        let mut t = Thermostat::new(28);
        t.warmer(10);
        assert_eq!(t.target(), MAX);
    }
}
     Running unittests src/lib.rs (target/debug/deps/heating-795830f7b1de3879)

running 2 tests
test tests::a_new_thermostat_keeps_its_target ... ok
test tests::warmer_never_passes_the_maximum ... ok

test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

You have read that summary line 46 times without needing it. Now it is yours, so take the five fields apart once: passed and failed are self-explanatory; ignored counts tests marked #[ignore], which Part 5 covers; measured is for nightly-only benchmarks and will always be 0 for you; and filtered out counts tests that exist but did not run because you passed a name filter. Note also that the test name is tests::a_new_thermostat… — the module path is part of the test's name, which is what makes filtering by module possible later.

Now the same run with a bug planted in capped, which drops the upper bound (degrees.clamp(MIN, MAX) becomes degrees.max(MIN)):

running 2 tests
test tests::a_new_thermostat_keeps_its_target ... ok
test tests::warmer_never_passes_the_maximum ... FAILED

failures:

---- tests::warmer_never_passes_the_maximum stdout ----

thread 'tests::warmer_never_passes_the_maximum' (192610) panicked at src/lib.rs:66:9:
assertion `left == right` failed
  left: 38
 right: 30
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace


failures:
    tests::warmer_never_passes_the_maximum

test result: FAILED. 1 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

error: test failed, to rerun pass `--lib`

Three sections, and each answers a different question. The per-test lines say which tests ran. The failures: block with the stdout capture says why each failure happened, and it is the only place the panic message appears. The short failures: list at the end is just names, so that with forty tests and six failures you can copy one name and re-run it alone. The final error: test failed, to rerun pass --lib is cargo telling you which target to narrow to — --lib for unit tests, --test <name> for one integration file.

Part 2 — Three macros, and what each failure tells you

You only need three, and the choice between them is entirely about what you want printed when the test fails. That is the whole design question: a passing test prints nothing interesting, so a macro earns its keep only by how much it tells you on the day it goes red.

MacroFails whenPrints
assert!(cond)cond is falsethe source text of cond
assert_eq!(a, b)a != bboth values, as left and right
assert_ne!(a, b)a == bboth values, the same way

Book: 11.1 — Testing equality with assert_eq! and assert_ne!

Prefer assert_eq! whenever you have an expected value to name, because a bare assert! throws away the numbers. Compare these two failures of the same bug — the room comparison in is_heating flipped to room > self.target. First, assert! on its own:

#[test]
fn a_cold_room_heats() {
    let t = Thermostat::new(20);
    assert!(t.is_heating(18));
}
thread 'tests::a_cold_room_heats' (196474) panicked at src/lib.rs:59:9:
assertion failed: t.is_heating(18)

That tells you the expression was false, which you could have guessed from the test's name. When the condition is a bool and there is nothing to compare, add the message yourself — every argument after the condition is handed to format!, so you can print whatever would have helped:

#[test]
fn a_cold_room_heats() {
    let t = Thermostat::new(20);
    assert!(
        t.is_heating(18),
        "a room at 18 must heat towards {}",
        t.target()
    );
}
thread 'tests::a_cold_room_heats' (196562) panicked at src/lib.rs:59:9:
a room at 18 must heat towards 20

Your shipped tests use this in a place worth copying. In tests/persist.rs the loop over corrupt lines ends with "line {:?} should be reported as a bad line", bad, because the assertion runs six times and the failure would otherwise not say which line broke it. That is the rule: if an assertion runs inside a loop, it needs a message naming the case, or a red test sends you back to guessing.

One requirement comes attached to assert_eq!, and you have already satisfied it without knowing. To print the two values, the macro needs Debug; to compare them, it needs PartialEq:

When the assertions fail, these macros print their arguments using debug formatting, which means the values being compared must implement the PartialEq and Debug traits.

Book: 11.1

This is why #[derive(Debug, PartialEq)] sits on Task, Status, Priority, Command, and Store — not decoration, a testing requirement. And it is why TaskError has that hand-written impl PartialEq from 0007: io::Error does not implement it, so the derive was impossible and you compared kind() instead. Every one of those impls exists so that assert_eq! can print something useful. Today you are finally the one calling it.

Part 3 — Two ways to test a failure

Code that works is the easy half. The interesting tests are the ones that pin down what happens when the input is wrong, and Rust gives you two tools because your code has two ways to fail: it panics, or it returns an Err.

For a panic, annotate the test with #[should_panic], and the test passes if and only if the body panics. Always give it expected, a substring of the panic message, or the test will happily pass on a panic that came from somewhere else entirely:

#[test]
#[should_panic(expected = "at most 30")]
fn refuses_a_high_target() {
    Thermostat::new(99);
}
running 2 tests
test tests::a_text_target_is_read_or_reported ... ok
test tests::refuses_a_high_target - should panic ... ok

Notice the - should panic marker in the result line: the runner tells you the test's polarity is inverted, which matters when you are reading someone else's suite. And here is the same test against a new whose upper-bound branch was given the lower bound's message by mistake — the code still panics on 99, but says the wrong thing:

thread 'tests::refuses_a_high_target' (197030) panicked at src/lib.rs:15:13:
target must be at least 5, got 99
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
note: panic did not contain expected string
      panic message: "target must be at least 5, got 99"
 expected substring: "at most 30"

Without expected that run would have been green, and the test would have been worthless: it would have proved only that something went wrong. The expected substring is what turns "it panicked" into "it panicked for the reason I meant".

For an Err, the tool is different and better suited to your crate: a test may return Result, which lets you use ? in its body. The test passes on Ok and fails on Err:

#[test]
fn a_text_target_is_read_or_reported() -> Result<(), String> {
    let mut t = Thermostat::new(20);
    t.set_from(" 21 ")?;              // an Err here fails the test
    assert_eq!(t.target(), 21);
    assert!(t.set_from("hot").is_err());
    Ok(())
}

Two rules come with that shape, and the second one is the trap. First, the return type has to be a Result whose error implements Debug — Result<(), TaskError> qualifies, since TaskError derives Debug. Second, from the book:

You can't use the #[should_panic] annotation on tests that use Result<T, E>. To assert that an operation returns an Err variant, don't use the question mark operator on the Result<T, E> value. Instead, use assert!(value.is_err()).

Book: 11.1 — Using Result<T, E> in tests

Read those two together and the division of labour is clear. Use ? for the steps that are merely setup — the save, the load, the completion that has to work before the interesting assertion can run — and use an explicit assert_eq!(…unwrap_err(), …) or assert!(…is_err()) for the failure you are actually testing. Your shipped tests/errors.rs does the second half already; the drill has you write the first.

Which means, honestly, that #[should_panic] has almost no place in your crate — and that is a result, not a gap. Since 0006 your code returns TaskError instead of panicking, so there is no panic left to pin. Learn the attribute because interview questions and other people's crates use it; reach for the Result form in your own.

Part 4 — Where tests live, and what each kind can see

Rust has exactly two homes for tests, and the choice is not stylistic — it decides what your test is allowed to touch:

Unit tests are small and more focused, testing one module in isolation at a time, and can test private interfaces. Integration tests are entirely external to your library and use your code in the same way any other external code would, using only the public interface and potentially exercising multiple modules per test.

Book: 11.3 — Test organization

A unit test lives in the same file as the code it tests, at the bottom, in a module with two attributes' worth of ceremony:

#[cfg(test)]          // compile this only for `cargo test`
mod tests {
    use super::*;     // pull the whole parent module into scope

    #[test]
    fn the_private_cap_holds_both_ends() {
        assert_eq!(capped(99), MAX);      // private fn, reachable
        assert_eq!(capped(-40), MIN);
        assert_eq!(capped(21), 21);
    }
}

Both lines earn their place. #[cfg(test)] means the module is not compiled into cargo build output at all, so tests cost nothing in the shipped binary. use super::* is what gives the test its reach: the tests module is an ordinary child module, and a child may see its parent's private items — which is the entire reason unit tests can test private functions. No annotation grants that privilege; the module tree does, exactly as chapter 7 described it.

An integration test lives in tests/, and gets a very different deal. Each file there is compiled as its own separate crate that uses yours from outside, so it sees precisely what a stranger on crates.io would see. Ask for anything private and the compiler says so — this is a real cargo test run of a tests/outside.rs that tries both:

error[E0603]: function `capped` is private
  --> tests/outside.rs:7:25
   |
 7 |     assert_eq!(heating::capped(99), 30);  // so is the helper
   |                         ^^^^^^ private function
   |
note: the function `capped` is defined here
  --> src/lib.rs:48:1
   |
48 | fn capped(degrees: i32) -> i32 {
   | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

error[E0616]: field `target` of struct `Thermostat` is private
 --> tests/outside.rs:6:18
  |
6 |     assert_eq!(t.target, 20);       // the field is private
  |                  ^^^^^^ private field
  |
help: a method `target` also exists, call it with parentheses
  |
6 |     assert_eq!(t.target(), 20);       // the field is private
  |                        ++

You have met E0616 before from the other side. In 0003 you made Store.tasks private and added the tasks() accessor, and every one of my 46 tests goes through that accessor because it has no choice. So the trade is now concrete: put a test in src/ and it can reach inside; put it in tests/ and it is forced to use the API you actually ship, which means it also notices when you break that API. Write both kinds, for different reasons — the file-local ones to pin down awkward internals, the external ones to pin down the contract.

Sharing a helper between two integration files has one gotcha, and it is worth spending a paragraph on because you will hit it in the drill. Since every file in tests/ is its own crate, a tests/common.rs full of helpers is compiled as a test crate of its own and shows up in the output as a pointless running 0 tests section. The fix is the older module-file spelling:

tests/
├── common/
│   └── mod.rs     ← helpers live here; not treated as a test crate
├── cli.rs         ← `mod common;` then `common::three_tasks()`
└── mine.rs        ← same, its own crate, its own copy

Book: 11.3 — Submodules in integration tests: “Files in subdirectories of the tests directory don't get compiled as separate crates or have sections in the test output.”

And now the rule this whole lesson turns on. It is one paragraph in the book, and it explains your stats bug exactly:

If our project is a binary crate that only contains a src/main.rs file and doesn't have a src/lib.rs file, we can't create integration tests in the tests directory and bring functions defined in the src/main.rs file into scope with a use statement. … This is one of the reasons Rust projects that provide a binary have a straightforward src/main.rs file that calls logic that lives in the src/lib.rs file.

Book: 11.3 — Integration tests for binary crates

Your crate has both files, which is why the tests can see Store at all. But your run function — the one that decides the order of the stats lines and whether clear says anything — is defined in main.rs, on the wrong side of that wall. No test can import it. The book's advice is the fix: main.rs should be small enough that it needs no test, and everything else belongs in the library. Step 3 of the drill moves run across.

Moving it is not enough on its own, though, and the second half is the more useful trick. A run that calls println! writes to the process's stdout, which a test cannot read. So instead of printing, take the destination as a parameter:

pub fn run(
    args: &[String],
    store: &mut Store,
    out: &mut impl Write,        // std::io::Write
) -> Result<(), TaskError>

main hands it io::stdout().lock() and behaves exactly as before. A test hands it a Vec<u8>, which implements Write, and then asserts on the bytes. That is the whole technique: a function that returns or writes its output can be tested; a function that prints its output cannot. It costs one parameter, and it is the single most reusable idea in this lesson — the same move makes an HTTP handler testable without a server, and it is the answer to the interview question “how would you test that?”

Part 5 — Running them: the flags worth knowing

Everything so far assumed a bare cargo test. Four flags cover the rest of daily use, and the first thing to know is where the separator goes: arguments before -- are read by cargo, arguments after it are read by the test binary cargo just built.

CommandWhat it does
cargo test warmerruns tests whose full name contains warmer
cargo test --libonly the unit tests inside src/
cargo test --test clionly tests/cli.rs
cargo test -- --show-outputalso print stdout from tests that passed
cargo test -- --ignoredonly the tests marked #[ignore]
cargo test -- --test-threads=1no parallelism

Book: 11.2 — Controlling how tests are run

Filtering matches on the whole test name, module path included, which is the payoff of that tests:: prefix from Part 1. A real run, with three of four tests filtered out:

$ cargo test warmer
     Running unittests src/lib.rs (target/debug/deps/heating-795830f7b1de3879)

running 1 test
test tests::warmer_never_passes_the_maximum ... ok

test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 3 filtered out; finished in 0.00s

Output capture is the behaviour that surprises people: a println! in a passing test is swallowed, and only reappears if the test fails. When you want to see it anyway, ask:

$ cargo test -- --show-output
running 1 test
test tests::warmer_never_passes_the_maximum ... ok

successes:

---- tests::warmer_never_passes_the_maximum stdout ----
target ended at 30


successes:
    tests::warmer_never_passes_the_maximum

test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

#[ignore] is for the test you want to keep but not run every time — the slow one, the one that needs a network. It takes a reason string, which the runner prints, and the ignored tests are still one command away:

#[test]
#[ignore = "slow: walks the whole range"]
fn every_legal_target_round_trips() {
    for degrees in MIN..=MAX {
        let mut t = Thermostat::new(MIN);
        t.set_from(&degrees.to_string()).expect("legal target");
        assert_eq!(t.target(), degrees);
    }
}
$ cargo test
running 4 tests
test tests::every_legal_target_round_trips ... ignored, slow: walks the whole range
test tests::refuses_a_high_target - should panic ... ok
test tests::the_private_cap_holds_both_ends ... ok
test tests::warmer_never_passes_the_maximum ... ok

test result: ok. 3 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 0.00s

$ cargo test -- --ignored
running 1 test
test tests::every_legal_target_round_trips ... ok

test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 3 filtered out; finished in 0.00s

The last flag comes with the one hard constraint of the whole chapter. Tests run in parallel, on threads, by default:

Because the tests are running at the same time, you must make sure your tests don't depend on each other or on any shared state, including a shared environment, such as the current working directory or environment variables.

Book: 11.2 — Running tests in parallel or consecutively

Your suite obeys this already, and now you can see why it was written that way. Every file-touching test calls temp_path(), which mixes the process id with an atomic counter to produce a path no other test will ever use. That is the first solution the book offers — one file per test. --test-threads=1 is the second, and it is a worse one: it is slower, and it hides the coupling instead of removing it. Reach for it to diagnose a flaky suite, not to fix one. The same reasoning is why no test of yours may set TASKS_FILE: environment variables are per-process, so a test that sets one is reaching into every other test running at that moment.

Part 6 — A test that cannot fail is not a test

Green tests are not evidence. Forty-six of them were green while stats printed its lines in the wrong order, and no amount of staring at the count would have told you. The only honest question about a test suite is: which bugs would it catch?

There is a mechanical way to ask it, called mutation testing. Plant a deliberate bug in a copy of the code, run the suite, and see whether it goes red. A bug the suite notices is killed. A bug it sleeps through survives, and every survivor is a precise, undeniable description of a missing test. This lesson ships six of them in 0009-mutants.sh: it copies your crate to a temp directory, applies one sed substitution, and runs your tests. Your own files are never touched.

Here is that script run against your crate exactly as it stands right now, before the drill:

$ bash ~/learn-rust/lessons/0009-mutants.sh ~/learn-rust/tasks
crate: /home/tan/learn-rust/tasks
  SKIP     stats-order   src/cli.rs does not exist yet
  SKIP     stats-zero    src/cli.rs does not exist yet
  SKIP     clear-count   src/cli.rs does not exist yet
  SURVIVED list-format   src/task.rs
  SURVIVED status-parse  src/task.rs
  SURVIVED command-case  src/command.rs

0 killed, 3 survived, 3 skipped

Read the six lines as a to-do list, because that is what they are. The three SKIPs are the mutations that live in src/cli.rs — the file you have not written yet, which is where run is going. The three SURVIVEDs are real bugs your 46 tests cannot see today: Display for Task could stop printing the priority, Status::parse could stop understanding in-progress, and Command::parse could stop accepting ADD in capitals, and every test would still pass. The drill's finishing condition is 6 killed, 0 survived.

One caveat, so you calibrate the tool correctly rather than worshipping it: a suite that kills every mutant is not a proven-correct suite, because my six mutants are not every possible bug. Mutation testing gives you a floor, not a ceiling. It is still the sharpest feedback available on a suite you just wrote, and it is far better than counting tests.

Check yourself before the drill

Six questions before you touch the keyboard. Answer each one out loud, in full sentences, before you reveal or click. Two of them revisit 0005–0008 rather than today's material, which is deliberate — retrieval of old work is what keeps it.

Tests

What actually makes a #[test] function fail?

Tests

A unit test and an integration test: where does each file live, and what can each one reach?

Tests

You write a test that returns Result<(), TaskError> and want to assert a call fails. What is the correct move?

Traits

Why can assert_eq! compare and print two Task values, and why did TaskError need a hand-written PartialEq?

Collections

Why can a test not assert on count_by_priority() by iterating the map and printing as it goes?

Modules

Your run lives in src/main.rs. Why can no file in tests/ import it, and what are the two changes that make its output testable?

The drill — 45 minutes, your own crate

Type it, do not paste it. The thermostat above is a different program. Keep the tests reference open — looking syntax up is free.

cd ~/learn-rust/tasks
cargo test                                   # 46 pass, as they did yesterday
bash ../lessons/0009-mutants.sh .            # 0 killed, 3 survived, 3 skipped

Those two lines are the starting position. Every test you write today is yours — I am shipping no new spec file, because the skill being built is writing the assertions rather than satisfying them. The finishing line is the mutant report reading 6 killed, 0 survived, 0 skipped, and all 46 existing tests still green.

Step 0 — the last 0008 leftover, one minute

Delete the commented-out for loop still sitting inside Store::find. The iterator version is one line above it and git remembers the old one.

Check: grep -c "for " src/store.rs prints 0.

Step 1 — your first #[test], in src/task.rs

Add a #[cfg(test)] mod tests at the bottom of src/task.rs with three tests, and one more at the bottom of src/stats.rs. All four target behaviour that none of my 46 tests touches — that is why they are worth your keystrokes rather than being duplicates.

The first pins the line format on disk: build a Task with a known id, title, priority and status, assert that to_line() produces exactly the string you expect, and assert that parsing that string back gives the task you started with. Both directions in one test, because a round trip that only goes one way proves nothing about the other.

The second pins every Status label, not just the two the CLI uses. Loop over the three variants, and for each one assert that Status::parse(status.label()) gives that variant back. Your in-progress arm is currently unreachable from the CLI and therefore completely untested — the loop covers it without you writing three near-identical tests. Give the assertion a failure message naming the label, per Part 2, or a red run will not say which variant broke.

The third pins what the user reads: the Display impl you wrote in 0005. Assert the exact line for a fresh task, then set its status to Done and assert the line again. Nothing in the suite has ever checked this string.

The fourth, in src/stats.rs, calls tally with a key that is owned rather than Copy — a closure returning String — and asserts both a count and the map's length. Your count_by_priority only ever hands tally a Copy key, so the generic function has never been exercised with anything else.

Check: cargo test --lib → 4 passed. Note the names in the output: task::tests::… and stats::tests::….

Forgotten what the test module looks like?

#[cfg(test)] then mod tests { then use super::*; — Part 4 has the whole shape. Without use super::* you get E0433: failed to resolve on the first type name, because the child module starts with an empty scope.

Step 2 — shared helpers, and tests that return Result

Create tests/common/mod.rs — the directory spelling from Part 4, not tests/common.rs — holding three helpers you will use from two files: temp_path() (copy the one from the top of tests/collections.rs; a helper worth sharing is a helper worth moving), args(&[&str]) -> Vec<String>, and three_tasks() -> Store which returns a store holding one task per priority with task 1 already completed. Put #![allow(dead_code)] at the top of the file: each test crate uses only some of the helpers, and without it the unused ones warn.

Then write tests/mine.rs with three tests, each returning Result<(), TaskError> so the setup steps can use ?:

Completing a task that is already done is not an error. Complete task 1 a second time and assert it is still Done. Your complete takes this path today; the test decides that the behaviour is deliberate rather than accidental, which is what a test is for.

A reload sees exactly what was saved. Take three_tasks(), clear the completed one, save to a temp_path(), load it back, and assert the loaded tasks equal the ones in memory. The shipped suite tests save-then-load, but never after a removal.

Saving a smaller store shortens the file. Save three tasks, remove the completed one, save again to the same path, then read the file with fs::read_to_string and assert it has two lines. If save ever stops truncating, this is the only test that will notice — and a save that appends instead of replacing is a data-loss bug, not a cosmetic one.

Check: cargo test --test mine → 3 passed, and a full cargo test shows no Running tests/common section. If you see one, you named the file tests/common.rs.

Step 3 — move run into the library

This is the structural step, and the point of it is Part 4's rule: nothing in main.rs can be tested, so almost nothing should live there.

Create src/cli.rs, declare it in lib.rs, and move run into it with this signature:

pub fn run(
    args: &[String],
    store: &mut Store,
    out: &mut impl Write,
) -> Result<(), TaskError>

Replace every println!(..) in the body with writeln!(out, ..)?. The ? is doing real work there: writeln! returns io::Result, and your From<io::Error> for TaskError from 0008's step 0 converts it — the second time that impl has paid for itself. Then main becomes: build the path, load the store, collect the args, take io::stdout().lock(), call run, save, and fail on either error.

While the code is open, fix the two defects from the top of this page. stats loops over [Priority::High, Priority::Medium, Priority::Low] — fully qualified, in that order — and prints each with writeln!(out, "{:<6} {}", p.label(), n)?, so the counts line up in a column. clear keeps the usize that remove_completed returns and prints cleared N completed. Those exact formats are what the tests in step 4 assert, and they are the output the 0008 session captured.

Check: grep -c "println!" src/cli.rs prints 0, and the CLI still behaves — from a scratch directory, with run(){ TASKS_FILE=t.txt cargo run -q --manifest-path ~/learn-rust/tasks/Cargo.toml -- "$@"; }:

$ run add "buy milk" high ; run add "call bank" ; run add "water plants" low
$ run done 1
$ run stats
high   1
medium 1
low    1
$ run clear
cleared 1 completed
$ run done 9 ; echo $?
error: no task with id 9
1

Step 4 — the tests that catch what 46 could not

Write tests/cli.rs. Start with a helper that runs one command against a store and gives back exactly what it printed — a Vec<u8> for out, then String::from_utf8:

fn output(command: &[&str], store: &mut Store) -> String {
    let mut out: Vec<u8> = Vec::new();
    run(&args(command), store, &mut out).expect("command must succeed");
    String::from_utf8(out).expect("output must be utf-8")
}

Then six tests, each asserting on the whole printed string with assert_eq! rather than searching it for a substring — an exact assertion is what kills the mutants, and the escaped \ns are part of the contract:

Check: cargo test --test cli → 6 passed. Full cargo test → 4 + 6 + 14 + 7 + 3 + 8 + 17 = 59 passed, of which 13 are yours.

Step 5 — hunt the mutants

Run the script against your crate. Every one of the six should now be reported killed:

$ bash ../lessons/0009-mutants.sh .
  killed   stats-order   src/cli.rs
  killed   stats-zero    src/cli.rs
  killed   clear-count   src/cli.rs
  killed   list-format   src/task.rs
  killed   status-parse  src/task.rs
  killed   command-case  src/command.rs

6 killed, 0 survived, 0 skipped

If one survives, do not adjust the script — read the mutation it names in the source of the script, work out which of your tests should have caught it, and fix that test. A survivor is never wrong: it is a bug that your suite genuinely cannot see. If one says SKIP, the pattern is not in your source, which usually means you spelled that line differently; the script prints the file so you can compare.

Then stop

Not today: doc tests (/// examples that run — chapter 14), #[bench], assert_cmd and predicates for testing the binary as a subprocess, proptest for generated inputs, and cargo-mutants, which is the real version of this lesson's script. Each is a small step from here, and none is on the path to the next gap.

What this closed

Chapter 11 moves to produced on the coverage map, and with it the last core gap in chapters 1–11. You have now written unit tests, integration tests, shared helpers, Result-returning tests, and an assertion on a command's exact output — plus the refactor that made the last one possible, which is the part an interviewer will actually probe.

What is left before the job-ready floor is short, and it is no longer about the book's core:

  1. ch 10.3 — lifetimes, as reading practice. You have now written two without noticing: titles_with returns Vec<&str> borrowed from &self, and Status::label returns &str borrowed from &self. Elision filled in both annotations for you, and reading the explicit form is a two-lesson job at most.
  2. serde, which replaces your to_line/FromStr pair with two derives — worth doing after writing them by hand, which you now have.
  3. Then axum, where the traits from 0005–0008 and the testing from today start paying rent together: a handler is just a function you can call from a test.

Take it outside

Here is a question with genuine disagreement behind it, which makes it a good one to ask people rather than docs. Your output helper asserts on the exact bytes a command prints, which means a wording change to cleared N completed breaks a test even though nothing is broken for the user. Some engineers call that a feature — the output is the contract, and changing it should be deliberate. Others call it a brittle test that will be deleted the first time it is inconvenient, and would assert only that the count appears somewhere in the line. Ask on users.rust-lang.org where they draw that line for CLI output, and what they do differently for output a machine parses versus output a human reads. The answers will teach you more about test design than any chapter, because it is a taste question and the book cannot have taste for you.

The five sentences worth keeping

  1. A test fails when its thread panics; every assertion macro is a wrapper that panics, so unwrap in a test is a legitimate assertion.
  2. Unit tests live beside the code in #[cfg(test)] mod tests and can see private items; integration tests live in tests/, are separate crates, and see only the public API.
  3. #[should_panic(expected = "…")] tests a panic, a -> Result<(), E> test lets you ? the setup, and the two cannot be combined.
  4. Nothing in src/main.rs is testable, so main stays thin and everything else moves to the library — and a function that writes to &mut impl Write is testable where one that calls println! is not.
  5. Tests run in parallel and share nothing safely, so give every file-touching test its own path; and judge a suite by the bugs it kills, never by the number of tests it contains.