Files, io::Error, and FromStr

Lesson 0007 · after 0006 · reading, then a 30-minute drill against a shipped test file · ~45 minutes

Every code block and every compiler message on this page was produced by running it today, in a scratch project. Nothing is written from memory. The demo domain is a weather log — your project is a task CLI, so nothing here pastes in. Translating is the work.

Where 0006 left you, and the one loose end

Twenty-four tests are green, and the error type behind them is real work: TaskError has seven variants, each with its own Display sentence, a source() that hands back the one wrapped cause, and an impl From<ParseIntError>. That is the full set of obligations for a std-compatible error type, and you wrote all of it from the signatures up.

There is one loose end, and it is worth looking at closely before adding anything new. The From impl you wrote is never actually called. Over in command.rs, the conversion is still done by hand — once in the done arm, once in remove:

let id: u32 = match id.parse() {
    Ok(n) => n,
    Err(e) => return Err(TaskError::BadId(e)),   // this IS From::from, typed out
};

Compare that with what From exists to do. Your impl says "given a ParseIntError, build a TaskError::BadId" — which is exactly what the Err arm above says, in five lines instead of zero. The impl is correct; it simply never gets reached, because match handles the error before ? would have had a chance to convert it. So the mechanism was learned and the reflex was not, which is the most common way a Rust concept half-lands.

Step 0 of today's drill deletes both blocks. Then today adds a second From impl, and this one you will not be able to route around by hand: the shape the code needs makes ? the only reasonable option, and the compiler stops the build until the impl exists.

Part 1 — A file is a String that can fail

A save file is, from Rust's point of view, nothing more exotic than a String that might not arrive. Two functions in std::fs cover everything a CLI of this size needs. Each one opens the file, does the work, and closes it again, all inside the single call — there is no handle to keep track of and nothing to remember to close:

fn read_to_string<P: AsRef<Path>>(path: P) -> io::Result<String>
fn write<P: AsRef<Path>, C: AsRef<[u8]>>(path: P, contents: C) -> io::Result<()>

std: fs::read_to_string · fs::write

The return types look unfamiliar, so take them apart before going further. io::Result<T> is not a new kind of Result; it is a type alias, declared in std as type Result<T> = std::result::Result<T, io::Error>. In other words the error half has already been filled in for you, because every function in that module fails the same way. Whenever you meet io::Result<String> in a signature, read it silently as Result<String, io::Error> and carry on — it is the same Result you have been matching on since chapter 9, and ? works on it exactly as you would expect.

(The P: AsRef<Path> in the signature is a convenience bound, and you can read past it for now. All it means is that you may pass a &str, a String, a &Path, or a PathBuf, and std will accept any of them. It is the same idea as the trait bounds from 0005, used to widen what a function will take.)

The interesting half is io::Error. Unlike ParseIntError, which really only means "that was not a number", an io::Error stands for dozens of distinct situations: the file does not exist, the path is a directory rather than a file, the process lacks permission to read it, the disk is full, the name is too long. All of those arrive as the same type, so the type alone cannot tell you what went wrong. You separate them by asking the value, using e.kind(), which returns a variant of the ErrorKind enum:

let missing = fs::read_to_string("/tmp/definitely-not-here.txt");
println!("{:?}", missing.map_err(|e| e.kind()));
missing -> Err(NotFound)

Keep that NotFound in mind. It looks like a small detail, but it turns into the most important design decision of the whole lesson, and Part 4 comes back to it: on the very first run of your CLI, the save file legitimately does not exist yet, and how you treat that one kind decides whether a fresh install works or looks broken.

Part 2 — The second From, and why it is legal

You asked, after the last lesson, whether a type can have two From impls. The answer decides how today's code is shaped, so here is the rule stated precisely: for any given source type T, there may be exactly one impl From<T> for YourType in the whole program. Write a second one for the same T and the compiler stops you with error[E0119]: conflicting implementations.

The reason is worth holding on to, because it explains a lot of Rust's trait rules. At the moment you write value?, the only information the compiler has is the pair of types involved: it is converting a ParseIntError into a TaskError. Nothing at that call site records what you meant by the failure. If two impls existed for that pair, there would be two possible answers and no way to choose between them, so the language forbids the ambiguity up front rather than picking one for you.

Nothing stops you from adding an impl for a different source type, though, and that is what today needs — one conversion for parse failures, and a new one for file failures:

impl From<ParseIntError> for TaskError { .. }   // T = ParseIntError   (0006)
impl From<io::Error>     for TaskError { .. }   // T = io::Error   (0007)

Write fs::write(path, out)? before the impl exists and the compiler tells you exactly what is missing. Real output:

error[E0277]: `?` couldn't convert the error to `TaskError`
  --> src/store.rs:82:29
   |
82 |         fs::write(path, out)?;
   |         --------------------^ the trait `From<std::io::Error>` is not implemented for `TaskError`
   |         |
   |         this can't be annotated with `?` because it has type `Result<_, std::io::Error>`
   |
note: `TaskError` needs to implement `From<std::io::Error>`
   = note: the question mark operation (`?`) implicitly performs a conversion
           on the error value using the `From` trait

The final note in that message is the part worth reading twice. ? has no special knowledge of io, and it is not a built-in shortcut for file handling: it simply calls From::from on whatever error it is given. It is the identical mechanism that has been quietly converting your ParseIntError into a BadId since 0006. Once you see ? as "return early, and run the error through From on the way out", every one of these messages becomes predictable rather than mysterious.

Part 3 — FromStr: the trait behind .parse()

You have been calling .parse() since the guessing game, and it has probably felt like a built-in piece of string handling. It is not. It is a trait method, and once you see the two declarations behind it, the whole thing stops being magic:

impl str {
    pub fn parse<F: FromStr>(&self) -> Result<F, F::Err> { .. }
}

pub trait FromStr: Sized {
    type Err;                                          // an ASSOCIATED TYPE
    fn from_str(s: &str) -> Result<Self, Self::Err>;
}

std: FromStr · Book: 20.2 — associated types

Read the two together. "42".parse::<u32>() works for exactly one reason: somewhere in std there is an impl FromStr for u32, and it declares type Err = ParseIntError. There is no special case for integers in the language. That means the door is open to you — implement the same trait for your own type and .parse() begins working on it immediately. This is a genuinely different move from writing a Task::from_line helper of your own. A helper is a function only your code knows about; implementing the trait means your type joins an interface that std and every other crate already speak, so any generic function taking F: FromStr will now accept a Task as well.

The unfamiliar line is type Err;, and it deserves its own paragraph because it is your first associated type. Think of it as a slot in the trait that the implementor fills in, once, and permanently — as opposed to a generic parameter, which the caller chooses at each call site. That distinction is exactly why ParseIntError appears nowhere in parse()'s signature: the signature says F::Err, meaning "whatever error type F declared when it implemented the trait". When you write type Err = TaskError in your impl, you are filling that slot for Task, and from then on line.parse::<Task>() is known to return Result<Task, TaskError> without anyone having to say so again.

Leave the slot out and the compiler is explicit about the missing piece:

error[E0046]: not all trait items implemented, missing: `Err`
   --> src/task.rs:103:1
    |
103 | impl FromStr for Task {
    | ^^^^^^^^^^^^^^^^^^^^^ missing `Err` in implementation
    |
    = help: implement the missing item: `type Err = /* Type */;`

And call .parse::<Task>() before implementing it at all:

error[E0277]: the trait bound `Task: FromStr` is not satisfied
  --> src/store.rs:98:29
   |
98 |             tasks.push(line.parse::<Task>()?);
   |                             ^^^^^ unsatisfied trait bound
   |
help: the trait `FromStr` is not implemented for `Task`

The whole thing, in the weather-log demo — run today, output below:

#[derive(Debug, PartialEq)]
struct Reading { station: String, celsius: f64 }

impl FromStr for Reading {
    type Err = String;                        // your error type goes in the slot

    fn from_str(line: &str) -> Result<Reading, String> {
        let (station, temp) =
            line.split_once('=').ok_or_else(|| line.to_string())?;
        Ok(Reading {
            station: station.to_string(),
            celsius: temp.parse().map_err(|_| line.to_string())?,
        })
    }
}
Ok(Reading { station: "oslo", celsius: -3.5 })
Ok(Reading { station: "lagos", celsius: 31.0 })
Err("broken")

One line in there is doing something you have not seen before, so look at the inner temp.parse().map_err(|_| ..)? closely. It parses a f64, and the failure it can produce is a ParseFloatError — but notice what that failure means in this context. It does not mean "the user typed a bad number at the keyboard"; it means "the line stored in this file is corrupt". Those are two different problems, they deserve two different variants, and only one of them can be the one that From produces automatically.

So the shape to remember is this: ? on its own handles the single canonical conversion, and map_err is how you name a different variant for any other meaning of the same error type. In the drill you will write both, a few lines apart, on the very same ParseIntError — the CLI path keeps ? and produces BadId, while the file path uses map_err and produces BadLine.

Part 4 — Three decisions the tests will hold you to

A missing file is not an error

Picture the very first time anyone runs your CLI. There is no tasks.txt yet, because nothing has ever created one. If load simply propagates whatever fs::read_to_string returns, that first run prints an error to stderr and exits 1 — the program looks broken before the user has done anything wrong. A missing file here is not a failure at all; it is the normal starting state, and it means "you have no tasks yet".

So this is one of the rare places where you deliberately catch a single kind of error and turn it into a successful result, while letting every other kind through untouched:

let text = match fs::read_to_string(path) {
    Ok(text) => text,
    Err(e) if e.kind() == io::ErrorKind::NotFound => return Ok(Store::new()),
    Err(e) => return Err(TaskError::Io(e)),      // permissions etc. still fail
};

The new syntax is Err(e) if .., which is called a match guard: an extra condition attached to an arm, so the arm only matches when the pattern fits and the condition holds. Here the first Err arm catches only NotFound, and anything else falls through to the arm below it.

It is worth being clear about what this is not, because the lazy version is tempting. This is not unwrap_or_default() and it is not .ok(). Both of those would treat every io failure as "no tasks" — so a permissions problem, or a disk that has gone read-only, would silently present the user with an empty list, and the next save would overwrite their real file with nothing. One specific kind of failure is expected; the rest genuinely are failures and must still be reported.

The saved format is not the Display format

You already have a Display for Task from lesson 0005, and it prints 1 [done] buy milk (high). That is a good sentence for a person reading a terminal, and it is a poor format to read back in: to reconstruct the task you would have to find the brackets and parentheses, while allowing for a title that might itself contain either. The format fights you because it was never designed to be parsed.

Storage has different requirements from presentation, so give it its own format — one with a separator that splits cleanly and a fixed field order:

1|done|high|buy milk

Two audiences, two formats: Display stays exactly as it is for the list command, and to_line is added beside it for the file. Do not be tempted to make one serve both.

Notice also that the title is placed last. That is deliberate, and it lets the title contain anything at all, including the separator itself. The reason is how splitn works: "a|b|c|d|e".splitn(4, '|') stops splitting after it has produced four pieces, so the fourth piece is the entire remainder, "d|e", with its | intact. Reach for split instead and a title containing a pipe silently loses everything after it. One of the shipped tests covers exactly this case.

#[derive(PartialEq)] will break, and that is informative

Add Io(io::Error) to the enum and the derive on line 3 fails:

error[E0369]: binary operation `==` cannot be applied to type `&std::io::Error`
  --> src/error.rs:12:8
   |
 3 | #[derive(Debug, PartialEq)]
   |                 --------- in this derive macro expansion
...
12 |     Io(std::io::Error),
   |        ^^^^^^^^^^^^^^
   |
note: `std::io::Error` does not implement `PartialEq`

The note at the bottom is the interesting part: io::Error deliberately does not implement PartialEq. That is a considered decision by the std authors, not an oversight — two io failures can carry the same message and still come from entirely different OS state, so "are these two errors equal?" has no honest answer. Your enum now contains one, and equality for the whole enum is therefore no longer derivable.

This is also a useful moment to see what derive actually is. It is not a language feature attached to the type; it is a code generator that writes an ordinary impl for you, comparing every field with ==. When one field cannot be compared, the generated line does not compile, and you get the error above pointing at the derive itself.

Dropping PartialEq is not an option, because the 0006 tests compare TaskError values with assert_eq! and you may not edit them. So write the impl by hand instead. This part is mechanical rather than conceptual — type it, understand the three notes underneath, and move on:

impl PartialEq for TaskError {
    fn eq(&self, other: &Self) -> bool {
        use TaskError::*;
        match (self, other) {
            (UnknownCommand(a), UnknownCommand(b))
            | (BadPriority(a), BadPriority(b))
            | (BadLine(a), BadLine(b)) => a == b,
            (BadId(a), BadId(b)) => a == b,
            (NotFound(a), NotFound(b)) => a == b,
            (Io(a), Io(b)) => a.kind() == b.kind(),  // the kind, not the error
            _ => std::mem::discriminant(self) == std::mem::discriminant(other),
        }
    }
}

Three pieces of that impl are new, and each is useful well beyond this one function.

First, match (self, other) matches on a tuple of two values at once. You build a temporary pair and pattern-match both halves together, which is how you ask "are these the same variant, and if so, are their payloads equal?" in a single expression.

Second, (A(a), A(b)) | (B(a), B(b)) => is an or-pattern: several patterns sharing one arm. Rust allows it here because every alternative binds the same names, a and b, at the same types, so the arm's body is valid whichever alternative matched. That is what lets three String-carrying variants share a single line instead of taking three.

Third, mem::discriminant returns an opaque value identifying which variant a value is, ignoring any payload. Comparing two of them answers "same variant?" without your having to name the variants at all, which handles the four payload-free cases in one line.

That last convenience has a real cost, and it is the kind of thing to notice now rather than discover later. The _ arm means the compiler will never again force you to update this impl when you add a variant — and a new variant carrying data would then be compared by variant alone, treating two different payloads as equal. It is an acceptable trade for a small error type, but it is a trade, not a free win.

Check yourself before the drill

Six questions before you touch the keyboard. Try to answer each one out loud, in full sentences, before you reveal or click — an answer you can say is an answer you have understood, and one you can only recognise on a page usually is not. Getting one wrong here costs you nothing; getting the same thing wrong twenty minutes into the drill costs you the drill.

Traits

You have impl From<ParseIntError> for TaskError. Which second impl is rejected by the compiler?

Traits

What is type Err in impl FromStr, and why is it not written impl FromStr<Err>?

Error handling

Your load calls fs::read_to_string(path)? and the file does not exist yet. What does the user see on their first ever run?

Error handling

Both done abc (a CLI argument) and a corrupt saved line produce a ParseIntError. You want two different variants. How, given only one From impl is allowed?

Ownership

fn load(path: &Path) -> Result<Store, TaskError> — no &self, and it returns a Store by value. Why is that not a copy, and where does the returned value live?

Enums

After loading two tasks from a file, why must Store recompute its next id instead of starting at 1?

The drill — 30 minutes, your own crate

Type it, do not paste it. The weather log above is a different program. Keep the files & FromStr reference open — looking syntax up is free.

cd ~/learn-rust/tasks
cp ../lessons/0007-persist-spec.rs tests/persist.rs
cargo test            # 8 new tests fail to compile — that is the starting line

Do not edit anything in tests/. All 24 existing tests must still pass. Target at the end: 32 passing.

Step 0 — pay off 0006 (2 minutes)

Delete both match id.parse() blocks in command.rs. Each becomes one line, and the From impl you wrote last lesson finally does its job:

let id: u32 = args.get(1).ok_or(TaskError::MissingId)?.parse()?;

Check: grep -c "match id.parse" src/command.rs prints 0, and cargo test --test errors still passes 7.

Step 1 — two new variants, and a hand-written PartialEq

Add to TaskError:

BadLine(String),    // a saved line that cannot be read back — carries the line
Io(io::Error),      // the file could not be read or written — wraps std's error

Then: remove PartialEq from the derive and write the impl from Part 4; add both Display arms; add Io(e) => Some(e) to source(); add From<io::Error>. Exact messages, checked by the tests:

Variantto_string() must be
BadLine("rubbish")cannot read saved line: rubbish
Io(..)cannot read or write the task file

Check: cargo build passes, and cargo test --test errors is still 7/7 — the old variants must behave exactly as before.

Step 2 — task.rs: one line out, one line in

Three additions:

Check: cargo test --test persist a_task_becomes and cargo test --test persist a_title_may both pass.

Stuck on repeating BadLine(line.to_string()) five times?

Bind it once as a closure and hand it to ok_or_else: let bad = || TaskError::BadLine(line.to_string()); then parts.next().ok_or_else(bad)?. Note ok_or_else, not ok_or — the first takes a closure and only builds the error when there is one, the second builds it every time. With a String allocation inside, that difference is real.

Step 3 — store.rs: save and load

pub fn save(&self, path: &Path) -> Result<(), TaskError>
pub fn load(path: &Path) -> Result<Store, TaskError>   // associated fn

save builds one String — one line per task, each ending in \n — and calls fs::write once. load reads, skips empty lines, parses each into a Task, and recomputes the next id. Both the missing-file guard and the id recomputation are in Part 4.

Check: cargo test --test persist passes all 8.

Step 4 — main.rs: load, act, save

Where the file lives should not be hard-coded into the logic. One line of ch12 gets you an override for free:

let path = PathBuf::from(
    env::var("TASKS_FILE").unwrap_or_else(|_| "tasks.txt".to_string()));

Then run takes the path, loads the store itself, and saves at the end. Note the ordering that falls out of ?: an error anywhere means save is never reached, so a failed command cannot corrupt the file.

Check — your CLI now remembers things:

$ cd $(mktemp -d)      # empty dir: proves a first run works with no file
$ run(){ TASKS_FILE=t.txt cargo run -q --manifest-path ~/learn-rust/tasks/Cargo.toml -- "$@"; }
$ run add "buy milk" high
added task 1
$ run add "call bank"
added task 2
$ run done 1
completed 1
$ run list
1 [done] buy milk (high)
2 [todo] call bank (medium)
$ cat t.txt
1|done|high|buy milk
2|todo|medium|call bank
$ echo "rubbish" >> t.txt ; run list ; echo $?
error: cannot read saved line: rubbish
1

That last one is the payoff for BadLine(String) carrying the line: the message names the exact text to go and fix. A String error, or a bare Io, could not.

Final check: cargo test → 17 + 7 + 8 = 32 passed. And add tasks.txt to .gitignore — it is user data, not source.

Then stop

Not today: serde and JSON (the real answer for storage, but it teaches a crate rather than a concept), file locking, and BufReader for files too big to hold in memory. Your task file is a few kilobytes; read_to_string is the correct tool, and reaching for a buffered reader here would be copying a pattern you do not need.

What this closed

Chapter 12 moves to produced on the coverage map: env::args, env::var, fs, stderr, and exit codes are now all in your own code. Two gaps left before the job-ready floor, and they are next:

  1. ch 8 + 13 — HashMap, map/filter/collect, plus your first hand-written generic function (lesson 0008). Your load loop is a collect::<Result<Vec<_>, _>>() waiting to happen — I left it as a for loop deliberately so 0008 has something of yours to rewrite.
  2. ch 11 — writing your own tests: you have now consumed 32 of mine and written zero (lesson 0009).

Take it outside

The forum ask from 0006 still stands and is now stronger: error.rs holds both user mistakes (UnknownCommand, BadPriority) and system failures (Io, BadLine) in one enum. Many Rust developers would split those into two types. Post it on users.rust-lang.org (Code Review category) and ask which they would do and why. That is a genuine design question with real disagreement behind it — the answer you get back is wisdom you cannot derive from the book.

The five sentences worth keeping

  1. io::Result<T> is just Result<T, io::Error>; e.kind() is how you tell one io failure from another.
  2. One impl From<T> for YourError per T — a different T is a new impl, a repeat T is E0119.
  3. ? for the canonical conversion, map_err for every other meaning of the same error.
  4. FromStr is what .parse() calls; type Err is an associated type — a slot the implementor fills, not a parameter the caller passes.
  5. A missing file on first run is expected, not exceptional: match ErrorKind::NotFound, and let every other kind fail loudly.