Files
learn-rust/lessons/0008-iterators-and-hashmap.html
T

785 lines
49 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<title>0008 — Iterators, HashMap, and your first generic function</title>
<link rel="stylesheet" href="../assets/style.css" />
</head>
<body>
<h1>Iterators, <code>HashMap</code>, and your first generic function</h1>
<p class="subtitle">Lesson 0008 · after <a href="0007-files-and-fromstr.html">0007</a> · reading, then a 35-minute drill against a shipped test file · ~50 minutes</p>
<div class="callout">
Every code block, every compiler message, and every terminal session on this page was produced by running it
today. Nothing is written from memory. The demo domain is a library shelf, defined in full in the next section —
your project is a task CLI, so nothing here pastes in. Translating is the work.
</div>
<h2>The demo domain, in full</h2>
<p>Every example on this page runs against the same four books, so it is worth reading the data model once
before the examples start. Then any snippet below can be read without guessing what a field is called or what
type it holds. This is the whole thing — two type definitions, two helpers, and one <code>Vec</code>:</p>
<pre><code>use Shelf::*; // so the examples can say Fiction instead of Shelf::Fiction
#[derive(Debug, PartialEq, Eq, Hash, Clone, Copy)]
enum Shelf { Fiction, History, Poetry } // fieldless, so Copy is free
#[derive(Debug)]
struct Book {
title: String, // owned text
shelf: Shelf, // which shelf it belongs on
borrowed: bool, // is it out on loan right now
}
fn book(title: &amp;str, shelf: Shelf, borrowed: bool) -&gt; Book {
Book { title: title.to_string(), shelf, borrowed }
}
// "fiction" -&gt; Some(Fiction), anything unknown -&gt; None
fn shelf_of(word: &amp;str) -&gt; Option&lt;Shelf&gt; {
match word {
"fiction" =&gt; Some(Fiction),
"history" =&gt; Some(History),
"poetry" =&gt; Some(Poetry),
_ =&gt; None,
}
}
// The data. Four books, three shelves, one of them out on loan.
let mut shelf: Vec&lt;Book&gt; = vec![
book("Dubliners", Fiction, true),
book("SPQR", History, false),
book("Ariel", Poetry, false),
book("Beloved", Fiction, false),
];
// A second value, used only where an element must be a plain string:
let titles: Vec&lt;&amp;str&gt; = vec!["Dubliners", "SPQR", "Ariel"];</code></pre>
<p>Two naming conventions to hold on to, because they are the only way to know an element's type at a glance.
<code>shelf</code> (lowercase) is the <code>Vec&lt;Book&gt;</code>, so <code>shelf.iter()</code> hands you
<code>&amp;Book</code> and the closure parameter is written <code>|b|</code>. <code>Shelf</code> (capitalised) is
the enum, so <code>b.shelf</code> is a field holding one of its three variants. And <code>titles</code> is a
<code>Vec&lt;&amp;str&gt;</code>, so <code>titles.iter()</code> hands you <code>&amp;&amp;str</code> and the closure
parameter is written <code>|t|</code>.</p>
<p>The mapping onto your own crate is exact, which is what makes the translation mechanical rather than
creative: <code>Book</code> is <code>Task</code>, <code>title</code> is <code>title</code>,
<code>Shelf</code> is <code>Priority</code>, and <code>borrowed</code> is <code>Status</code>. So when a
snippet below counts books per shelf, you are reading the <code>count_by_priority</code> you are about to
write.</p>
<h2>Where 0007 left you, and one prediction I got wrong</h2>
<p>Thirty-two tests are green and the persistence layer behind them is real: <code>fs::read_to_string</code> and
<code>fs::write</code>, a <code>NotFound</code> match guard so a first run does not look broken, an
<code>impl FromStr for Task</code> with its own associated <code>type Err</code>, and a hand-written
<code>PartialEq</code> because <code>io::Error</code> refuses to have one. The best signal is in
<code>command.rs</code>: both hand-rolled <code>match id.parse()</code> blocks are gone, replaced by
<code>.parse()?</code>. That was the whole point of step 0, and it landed — <code>?</code> plus <code>From</code>
is now a reflex rather than a fact.</p>
<p>But I predicted something in 0007 that turned out to be false, and it is worth a paragraph because the lesson
generalises. I claimed the compiler would <em>force</em> you to write <code>impl From&lt;io::Error&gt; for
TaskError</code>, because <code>?</code> on <code>fs::write</code> would be the only reasonable shape. You wrote
this instead:</p>
<pre><code>fs::write(path, contents).map_err(TaskError::Io)</code></pre>
<p>That is perfectly good Rust. <code>TaskError::Io</code> is a tuple-variant constructor, which means it is also
a function of type <code>fn(io::Error) -&gt; TaskError</code>, so handing it straight to <code>map_err</code> is
idiomatic and allocation-free. No <code>From</code> impl needed, no error, and the test still passes. My claim
was simply wrong: <strong>a compiler error can only force a design when no legal alternative exists</strong>, and
here a legal alternative existed.</p>
<p>So which one should you write? Both are correct, and the difference is leverage rather than style.
<code>map_err</code> converts at one call site; <code>From</code> converts at <em>every</em> call site,
including ones you have not written yet, and it is what makes bare <code>?</code> work on any function in
<code>std::fs</code>, <code>std::io</code>, or a future crate that returns an <code>io::Error</code>. Today's
<code>save</code> gets rewritten anyway, and the rewrite is shorter when <code>?</code> just works — so step 0 of
the drill writes the impl and deletes the <code>map_err</code>. One line each way.</p>
<h2>Part 1 — An iterator is a lazy machine with one button</h2>
<p>You have written iterator code already, in bursts: <code>env::args().skip(1).collect()</code> in
<code>main</code>, <code>self.tasks.iter().find(|t| t.id == id)</code> in <code>find</code>, and
<code>.map(|t| t.id).max().unwrap_or(0)</code> in <code>load</code>. What you have not had yet is the model
underneath them, so each one was memorised separately. The model is unusually small — one trait, one method:</p>
<pre><code>pub trait Iterator {
type Item;
fn next(&amp;mut self) -&gt; Option&lt;Self::Item&gt;;
// ~75 more methods, all with default bodies built on next()
}</code></pre>
<p class="cite">Book: <a href="https://doc.rust-lang.org/stable/book/ch13-02-iterators.html">13.2 — Processing a
series of items with iterators</a> · std:
<a href="https://doc.rust-lang.org/std/iter/trait.Iterator.html">Iterator</a></p>
<p>That is the entire interface. <code>next</code> hands back <code>Some(item)</code> until the sequence runs
out, then <code>None</code> forever. Notice <code>type Item</code>: it is an associated type, the same
mechanism you filled in as <code>type Err</code> when you implemented <code>FromStr</code> last lesson. Every
other method — <code>map</code>, <code>filter</code>, <code>find</code>, <code>collect</code>, <code>sum</code>
— is a default method written in terms of <code>next</code>. Which is why learning the vocabulary is cheap:
there is no new machinery behind any of them, only different ways of pressing the same button.</p>
<p>The one property that trips everybody up is laziness. The book states it flatly:</p>
<blockquote>
<p>In Rust, iterators are <em>lazy</em>, meaning they have no effect until you call methods that consume the
iterator to use it up.</p>
</blockquote>
<p class="cite">Book: <a href="https://doc.rust-lang.org/stable/book/ch13-02-iterators.html">13.2</a></p>
<p>Building a chain of adapters does no work and touches no elements. It only describes work. Write a chain and
forget to finish it, and the closure never runs even once — the compiler warns, because the warning is the only
thing standing between you and a silently dead line of code:</p>
<pre><code>titles.iter().map(|t| t.to_uppercase());</code></pre>
<pre><code>warning: unused `Map` that must be used
--&gt; examples/e1.rs:3:5
|
3 | titles.iter().map(|t| t.to_uppercase());
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
= note: iterators are lazy and do nothing unless consumed
= note: `#[warn(unused_must_use)]` (part of `#[warn(unused)]`) on by default</code></pre>
<p>So every chain has exactly two parts, and it is worth naming them because the names tell you where a chain
must end. <strong>Adapters</strong> take an iterator and return another iterator: <code>map</code>,
<code>filter</code>, <code>enumerate</code>, <code>skip</code>, <code>take</code>, <code>rev</code>. They are
lazy, and they compose. <strong>Consumers</strong> take an iterator and return something that is not an
iterator: <code>collect</code>, <code>find</code>, <code>position</code>, <code>count</code>, <code>sum</code>,
<code>max</code>, <code>any</code>, <code>for_each</code>. They do the work, and a chain that does not end in
one has not run.</p>
<p>The performance question answers itself once you see the structure, and it matters for the job you are aiming
at. A chain of adapters is not a chain of temporary vectors — each adapter is a small struct wrapping the previous
one, and the whole tower compiles down to a single pass. That is why rewriting a <code>for</code> loop as an
iterator chain costs nothing at runtime, and why nobody in Rust treats the choice as a speed trade-off.</p>
<h2>Part 2 — Three ways to iterate, and the ownership behind each</h2>
<p>Before any adapter runs you have to say <em>how</em> you want the elements, and this is the one place where
iterators meet the borrow rules. There are three methods, they differ only in the ownership they hand out, and
picking the wrong one is the most common way an iterator chain fails to compile:</p>
<table>
<tr><th>Call</th><th>Item type</th><th>Use it when</th></tr>
<tr><td><code>v.iter()</code></td><td><code>&amp;T</code></td><td>You are reading. The collection survives.</td></tr>
<tr><td><code>v.iter_mut()</code></td><td><code>&amp;mut T</code></td><td>You are editing in place. It survives.</td></tr>
<tr><td><code>v.into_iter()</code></td><td><code>T</code></td><td>You want the elements out. It is consumed.</td></tr>
</table>
<p class="cite">Book: <a href="https://doc.rust-lang.org/stable/book/ch13-02-iterators.html">13.2 — “if we want
to create an iterator that takes ownership … we can call <code>into_iter</code>”</a></p>
<p>This is the same three-way choice you already make with <code>&amp;self</code>, <code>&amp;mut self</code>,
and <code>self</code> in a method signature, applied one element at a time — so nothing new is being introduced,
only a new place for a rule you already know. It also explains a piece of your own code you may have written
without reading: your <code>complete</code> uses <code>iter_mut</code> because it assigns to
<code>task.status</code>, while <code>find</code> uses <code>iter</code> because it only looks. Swap them and
neither compiles.</p>
<p>One trap deserves seeing before you hit it in the drill, because the error message is about a borrow and the
cause is an iterator. When you keep the result of an <code>iter_mut</code> chain in a variable, the mutable
borrow of the whole collection stays alive for as long as that variable does:</p>
<pre><code>struct Library { books: Vec&lt;String&gt; } // titles only, to keep the error bare
impl Library {
fn rename(&amp;mut self, from: &amp;str, to: &amp;str) {
let book = self.books.iter_mut().find(|b| *b == from).unwrap();
println!("{} books", self.books.len()); // asks for a second borrow
*book = to.to_string();
}
}</code></pre>
<pre><code>error[E0502]: cannot borrow `self.books` as immutable because it is also borrowed as mutable
--&gt; examples/e4.rs:5:44
|
4 | let book = self.books.iter_mut().find(|b| *b == from).unwrap();
| ---------- mutable borrow occurs here
5 | println!("renaming 1 of {} books", self.books.len());
| ^^^^^^^^^^ immutable borrow occurs here
6 | *book = to.to_string();
| ----- mutable borrow later used here</code></pre>
<p>Read the three annotations as a timeline and the rule falls out: the borrow begins at
<code>iter_mut()</code>, and it ends after the <em>last use</em> of <code>book</code>, not at the end of the
statement that created it. Anything else touching <code>self.books</code> in between is a second borrow, and
that is exactly the rule from chapter 4. The fix is to reorder — read the length first, or finish with
<code>book</code> before asking. Nothing about iterators is special here; they just make the overlap easy to
write by accident.</p>
<h2>Part 3 — The verbs you will use every day</h2>
<p>Here is the whole working vocabulary, run against a shelf of books. Read the calls beside their real output
rather than trying to memorise signatures — the shapes are what you want in your fingers:</p>
<pre><code>// the same four books from the top of the page
let mut shelf: Vec&lt;Book&gt; = vec![
book("Dubliners", Fiction, true), book("SPQR", History, false),
book("Ariel", Poetry, false), book("Beloved", Fiction, false),
];
let all_titles: Vec&lt;&amp;str&gt; = shelf.iter().map(|b| b.title.as_str()).collect();
shelf.iter_mut().for_each(|b| b.borrowed = false);
let fiction: Vec&lt;&amp;str&gt; = shelf.iter()
.filter(|b| b.shelf == Shelf::Fiction)
.map(|b| b.title.as_str())
.collect();
shelf.iter().find(|b| b.title == "Ariel").map(|b| b.shelf);
shelf.iter().position(|b| b.title == "Ariel");
shelf.iter().any(|b| b.borrowed); // bool
shelf.iter().filter(|b| !b.borrowed).count(); // usize
shelf.retain(|b| b.shelf != Shelf::Poetry); // delete in place</code></pre>
<pre><code>all_titles -&gt; ["Dubliners", "SPQR", "Ariel", "Beloved"]
iter_mut -&gt; every book returned, borrowed set to false on each
fiction -&gt; ["Dubliners", "Beloved"]
find -&gt; Some(Poetry) // the ITEM, mapped: Option&lt;Shelf&gt;
position -&gt; Some(2) // the INDEX: Option&lt;usize&gt;, Ariel is 3rd
any / count -&gt; false / 4 // nothing is borrowed now, so all 4 are in
retain -&gt; removed 1, 3 left // Ariel was the only poetry book</code></pre>
<p>Two of those are worth a second look, because they are the ones your own <code>store.rs</code> currently
writes out longhand as loops.</p>
<p><strong><code>find</code> versus <code>position</code></strong> is a question about what you need next.
<code>find</code> gives you the element, which is what <code>complete</code> wants — it has to assign to
<code>status</code>. <code>position</code> gives you the index, which is what <code>remove</code> wants —
<code>Vec::remove</code> takes an index, not an element. Both return an <code>Option</code>, and both compose
straight into the error handling you already have: <code>.ok_or(TaskError::NotFound(id))?</code> turns a
<code>None</code> into your own error and unwraps the rest, collapsing a nine-line loop into one expression.</p>
<p><strong><code>retain</code></strong> is the one that saves you from a genuine bug. Deleting several elements
from a <code>Vec</code> by index in a loop is a classic error: each removal shifts everything after it down one,
so the loop skips elements. <code>retain</code> takes a predicate meaning “keep this one” and does a single
compacting pass. It is a method on <code>Vec</code> rather than on <code>Iterator</code>, because it mutates
the collection in place — one of several useful methods that live on the collection rather than on the trait.</p>
<p class="cite">std: <a href="https://doc.rust-lang.org/std/vec/struct.Vec.html#method.retain">Vec::retain</a>
· <a href="https://doc.rust-lang.org/std/iter/trait.Iterator.html#method.position">Iterator::position</a></p>
<h2>Part 4 — <code>collect</code> is the interesting one</h2>
<p><code>collect</code> looks like “make a <code>Vec</code>”, and that undersells it enough to hide the single
most useful trick in this lesson. Its real signature says something much stronger:</p>
<pre><code>fn collect&lt;B: FromIterator&lt;Self::Item&gt;&gt;(self) -&gt; B</code></pre>
<p class="cite">std: <a href="https://doc.rust-lang.org/std/iter/trait.Iterator.html#method.collect">Iterator::collect</a>
· <a href="https://doc.rust-lang.org/std/iter/trait.FromIterator.html">FromIterator</a></p>
<p>Read that as: <em>collect will build any type that knows how to be built from this kind of item</em>. The
target is chosen by <code>B</code>, and <code>B</code> is decided by you, at the call site — which is why
<code>collect</code> is the one method where you routinely have to state a type. Leave it out and the compiler
has nothing to go on:</p>
<pre><code>let shouted = titles.iter().map(|t| t.to_uppercase()).collect();</code></pre>
<pre><code>error[E0283]: type annotations needed
--&gt; examples/e2.rs:3:9
|
3 | let shouted = titles.iter().map(|t| t.to_uppercase()).collect();
| ^^^^^^^ ------- type must be known at this point
|
= note: multiple `impl`s satisfying `_: FromIterator&lt;String&gt;` found in the `alloc` crate:
- impl FromIterator&lt;String&gt; for Box&lt;str&gt;;
- impl FromIterator&lt;String&gt; for String;</code></pre>
<p>The note is the teaching. This is not the compiler being fussy about vectors — it is telling you that several
types can be built from a stream of <code>String</code>s and it will not guess which one you meant. You answer
either on the left, <code>let shouted: Vec&lt;String&gt; = ...</code>, or on the right with a turbofish,
<code>.collect::&lt;Vec&lt;String&gt;&gt;()</code>. Both are common; pick whichever reads better in the line.</p>
<p>Now the trick. <code>Result</code> and <code>Option</code> both implement <code>FromIterator</code>, so
<strong>an iterator of <code>Result</code>s can collect into a single <code>Result</code> holding a
<code>Vec</code></strong>. The same chain, with only the target type changed, gives two entirely different
answers:</p>
<pre><code>let words = ["fiction", "history", "rubbish", "poetry"];
let each: Vec&lt;Option&lt;Shelf&gt;&gt; = words.iter().map(|w| shelf_of(w)).collect();
let all: Option&lt;Vec&lt;Shelf&gt;&gt; = words.iter().map(|w| shelf_of(w)).collect();
let numbers: Result&lt;Vec&lt;u32&gt;, _&gt; =
"1 2 x 4".split(' ').map(str::parse::&lt;u32&gt;).collect();</code></pre>
<pre><code>Vec&lt;Option&gt; -&gt; [Some(Fiction), Some(History), None, Some(Poetry)]
Option&lt;Vec&gt; -&gt; None
Result&lt;Vec&gt; -&gt; Err("invalid digit found in string")</code></pre>
<p class="cite">std: <a href="https://doc.rust-lang.org/std/result/enum.Result.html#impl-FromIterator%3CResult%3CA,+E%3E%3E-for-Result%3CV,+E%3E">impl
FromIterator&lt;Result&lt;A, E&gt;&gt; for Result&lt;V, E&gt;</a></p>
<p><code>Vec&lt;Option&lt;Shelf&gt;&gt;</code> keeps every outcome, hole included. <code>Option&lt;Vec&lt;Shelf&gt;&gt;</code>
means all-or-nothing: the first <code>None</code> ends the iteration and the whole result is <code>None</code>.
That short-circuit is not a detail — it is the reason this is the right tool for reading a file. Your
<code>load</code> currently loops, parses each line, and pushes into a <code>Vec</code>, with
<code>?</code> inside the loop. One <code>collect</code> replaces all of it:</p>
<pre><code>let tasks: Vec&lt;Task&gt; = contents.lines()
.map(str::parse)
.collect::&lt;Result&lt;Vec&lt;Task&gt;, TaskError&gt;&gt;()?;</code></pre>
<p>The behaviour is exactly what a save file wants. Every line parses and you get the tasks; one line is corrupt
and you get that line's error and nothing else — no half-loaded store to accidentally save back over the good
file. And notice <code>map(str::parse)</code>: you can pass a function <em>path</em> where a closure is expected,
because <code>|line| line.parse()</code> and <code>str::parse</code> are the same function. Which
<code>parse</code>, of the many possible, is settled by the collect target — the <code>Vec&lt;Task&gt;</code> tells
the compiler to look for <code>Task</code>'s <code>FromStr</code> impl, the one you wrote last lesson.</p>
<p>One line-splitting detail comes with this rewrite, and it is a real trap rather than trivia:</p>
<pre><code>let file = "1|fiction|Dubliners\n2|poetry|Ariel\n";
file.split('\n') // -&gt; ["1|fiction|Dubliners", "2|poetry|Ariel", ""]
file.lines() // -&gt; ["1|fiction|Dubliners", "2|poetry|Ariel"]</code></pre>
<p class="cite">std: <a href="https://doc.rust-lang.org/std/primitive.str.html#method.lines">str::lines</a></p>
<p><code>split('\n')</code> yields an empty final piece for a file that ends in a newline, because the text
after the last separator is the empty string. That is why your current <code>load</code> needs an
<code>if !items.is_empty()</code> guard — the guard exists to paper over the wrong splitter.
<a href="https://doc.rust-lang.org/std/primitive.str.html#method.lines"><code>lines()</code></a> is built for
this job: it treats the trailing newline as a terminator rather than a separator, and it strips a
<code>\r\n</code> too, which is free Windows compatibility. Switch splitters and the guard disappears — after
which a blank line in the <em>middle</em> of a file is no longer silently skipped but reported as a bad line,
which is the honest answer for a corrupt file. A shipped test pins that behaviour.</p>
<h2>Part 5 — <code>HashMap</code>, and the two traits a key must have</h2>
<p><code>HashMap&lt;K, V&gt;</code> is the last of the three common collections, and it is the one you have not
used at all. It is a lookup by key rather than by position, it lives on the heap like <code>Vec</code>, and it
is not in the prelude, so it needs an import:</p>
<pre><code>use std::collections::HashMap;
let mut counts: HashMap&lt;Shelf, usize&gt; = HashMap::new();
counts.insert(Shelf::Fiction, 2);
counts.get(&amp;Shelf::Fiction); // Option&lt;&amp;usize&gt; — may be absent
counts.get(&amp;Shelf::Poetry).copied().unwrap_or(0); // absent counts as 0
for (shelf, n) in &amp;counts { } // ARBITRARY order — never trust it</code></pre>
<p class="cite">Book: <a href="https://doc.rust-lang.org/stable/book/ch08-03-hash-maps.html">8.3 — Storing keys
with associated values in hash maps</a></p>
<p>Two things there are easy to skim past and expensive to learn later. <code>get</code> returns an
<code>Option&lt;&amp;V&gt;</code>, so “missing key” is a value you handle rather than a crash — the same shape as
<code>Vec::get</code>. And iteration order is arbitrary and not stable between runs. If a user is going to read
your output, you must impose an order yourself; the <code>stats</code> command in today's drill prints high,
medium, low in a fixed sequence for exactly that reason.</p>
<p>The idiom that makes hash maps worth their weight is <code>entry</code>. Counting things is the standard
example, and the book's version is four lines:</p>
<pre><code>for b in &amp;shelf {
*counts.entry(b.shelf).or_insert(0) += 1;
}</code></pre>
<pre><code>entry() -&gt; {Fiction: 2, History: 1}</code></pre>
<p class="cite">Book: <a href="https://doc.rust-lang.org/stable/book/ch08-03-hash-maps.html">8.3 — Listing
8-25, counting occurrences of words</a></p>
<p>Take that line apart slowly, because it is dense and it is everywhere in real Rust.
<code>entry(key)</code> returns an <code>Entry</code>, an enum standing for a slot that may or may not be
filled. <code>or_insert(0)</code> fills it with <code>0</code> if it was empty, and either way hands back a
<code>&amp;mut usize</code> pointing into the map. <code>*</code> follows that reference so
<code>+= 1</code> lands on the number itself. The whole thing is one hash lookup — the version you would write
by hand, <code>if !map.contains_key(k) { map.insert(k, 0) }</code> followed by a <code>get_mut</code>, costs
two or three and reads worse.</p>
<p>Now the part that is specific to Rust. A key type must implement <code>Eq</code> and <code>Hash</code>, and
you will meet that rule at a call site from today's drill — so here is the line the next two errors point at,
before they point at it. <code>Store::count_by_priority</code> is one line long, and it hands the work to a
small generic function called <code>tally</code>. You write both in the drill; Part 6 builds
<code>tally</code> from this signature:</p>
<pre><code>// src/stats.rs — Part 6 explains it and the drill writes the body
pub fn tally&lt;T, K, F&gt;(items: &amp;[T], key: F) -&gt; HashMap&lt;K, usize&gt;
// src/store.rs:68 — count_by_priority, in full
tally(self.tasks(), |task| task.priority)</code></pre>
<p>Read that as “count the items, grouped by whatever the closure pulls out of each one”. It is all you need
for the errors below; the three type parameters and the body are Part 6's job. Your <code>Priority</code>
implements neither <code>Eq</code> nor <code>Hash</code>, so the first attempt at using it as a key fails twice
over:</p>
<pre><code>error[E0277]: the trait bound `Priority: Eq` is not satisfied
--&gt; src/store.rs:68:9
|
68 | tally(self.tasks(), |task| task.priority)
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ the trait `Eq` is not implemented for `Priority`
|
help: consider annotating `Priority` with `#[derive(Eq)]`
error[E0277]: the trait bound `Priority: Hash` is not satisfied
help: consider annotating `Priority` with `#[derive(Hash)]`</code></pre>
<p>The requirement is not bureaucracy; it is the data structure stating its contract. To find a key the map
hashes it to pick a bucket, then compares for equality inside that bucket — so a key it cannot hash or cannot
compare is a key it cannot store. <code>Hash</code> gives it the first, <code>Eq</code> the second.</p>
<p><code>Eq</code> is worth understanding rather than just deriving, since you already have
<code>PartialEq</code> and this looks like a duplicate. It is not: <code>Eq</code> is a marker with no methods
of its own, and it promises one extra property that <code>PartialEq</code> does not — that every value equals
itself. The famous exception is <code>f64</code>, where <code>NAN != NAN</code>, which is precisely why
<code>f64</code> implements <code>PartialEq</code> but not <code>Eq</code>, and why a <code>f64</code> cannot
be a <code>HashMap</code> key. A three-variant enum has no such problem, so the derive is honest.</p>
<p class="cite">std: <a href="https://doc.rust-lang.org/std/cmp/trait.Eq.html">Eq</a> ·
<a href="https://doc.rust-lang.org/std/hash/trait.Hash.html">Hash</a></p>
<p>Fix those two and a third error appears, which is the most instructive of the set:</p>
<pre><code>error[E0507]: cannot move out of `task.priority` which is behind a shared reference
--&gt; src/store.rs:68:36
|
68 | tally(self.tasks(), |task| task.priority)
| ^^^^^^^^^^^^^ move occurs because `task.priority` has type `Priority`,
| which does not implement the `Copy` trait
|
note: if `Priority` implemented `Clone`, you could clone the value</code></pre>
<p>The closure receives <code>&amp;Task</code> — a borrow — and a map key has to be owned, since the map keeps
it. Reading <code>task.priority</code> out of a borrow is a move out of something you do not own, which is
E0507, one of the most common errors in real Rust. Three fixes exist and they are not equivalent:
<code>.clone()</code> works but is noise for three variants; <code>#[derive(Clone, Copy)]</code> makes
<code>Priority</code> behave like <code>u32</code>, copied implicitly wherever it is read; keying by
<code>task.priority.label()</code> sidesteps it by using a <code>&amp;str</code> instead. Derive
<code>Copy</code>. A fieldless enum is a single small integer at runtime, copying it is free, and it is what std
does for its own small enums such as <code>ErrorKind</code>.</p>
<h2>Part 6 — Your first generic function</h2>
<p>Counting tasks by priority is one specific job, and today you will write it once and never again — because the
function you write is generic over what it counts. This is your first hand-written generic, so here it is whole,
and then taken apart:</p>
<pre><code>use std::collections::HashMap;
use std::hash::Hash;
pub fn tally&lt;T, K, F&gt;(items: &amp;[T], key: F) -&gt; HashMap&lt;K, usize&gt;
where
K: Eq + Hash,
F: Fn(&amp;T) -&gt; K,
{
let mut counts = HashMap::new();
for item in items {
*counts.entry(key(item)).or_insert(0) += 1;
}
counts
}</code></pre>
<p>Three type parameters, and each one is there for a reason. <code>T</code> is the element type, and the
function never looks inside a <code>T</code>, which is exactly why it works on tasks and on strings alike.
<code>K</code> is the key type, and it carries the bound <code>Eq + Hash</code> — not because <code>tally</code>
cares, but because the <code>HashMap</code> it returns does. <code>F</code> is the closure type. Every closure in
Rust has its own anonymous type, so the only way to accept one is a type parameter bounded by
<code>Fn(&amp;T) -&gt; K</code>, which reads as “anything callable that takes a <code>&amp;T</code> and returns a
<code>K</code>”.</p>
<p>The <code>where</code> clause is worth seeing as the point of the exercise rather than syntax to tolerate. It
is a contract in both directions: callers must supply types that satisfy it, and inside the body you may use
exactly the operations it guarantees and nothing else. That is why generics in Rust do not blow up at the call
site the way C++ templates can — the bounds are checked once, against the definition. Try to call
<code>item.to_string()</code> in there and it will not compile, because nothing in the clause promised
<code>T: Display</code>.</p>
<p>The payoff is that one definition serves cases that have nothing to do with each other:</p>
<pre><code>tally(&amp;shelf, |b| b.shelf) // -&gt; {History: 1, Fiction: 2}
tally(&amp;["a", "bb", "cc"], |w| w.len()) // -&gt; {1: 1, 2: 2}</code></pre>
<p>Two calls, two different <code>T</code>, two different <code>K</code>, and — this is the part that matters
for the interviews you are aiming at — no runtime cost for the generality. Rust monomorphises: it compiles one
specialised copy of <code>tally</code> per combination of types actually used, so each call site gets code as
tight as if you had written that version by hand.</p>
<p class="cite">Book: <a href="https://doc.rust-lang.org/stable/book/ch10-01-syntax.html">10.1 — Generic data
types</a> and <a href="https://doc.rust-lang.org/stable/book/ch13-01-closures.html">13.1 — Closures</a></p>
<p>Last note before the drill, and it is a taste question rather than a rule. <code>tally</code>'s body keeps a
<code>for</code> loop, on purpose. Iterators replace loops that <em>search</em>, <em>transform</em>, or
<em>collect</em> — those have a named adapter and the chain reads better than the loop. A loop that folds many
items into one accumulator is the case where a loop is still the clearest thing to write; the iterator version
exists, <code>fold</code>, and here it would be harder to read for no gain. The drill's grep check is scoped to
<code>store.rs</code> for exactly this reason.</p>
<h2>Check yourself before the drill</h2>
<p>Six questions before you touch the keyboard. Answer each one out loud, in full sentences, before you reveal or
click. An answer you can say is an answer you have understood; one you can only recognise on the page usually is
not. Getting one wrong here costs nothing — getting it wrong twenty minutes into the drill costs you the drill.</p>
<div class="q" data-type="mcq" data-topic="Iterators">
<p class="topic">Iterators</p>
<p class="prompt">How many times does the closure run in <code>v.iter().map(|x| f(x));</code> — with no <code>collect</code>?</p>
<div class="options">
<button class="opt" data-correct="true">Not once, and it warns</button>
<button class="opt" data-correct="false">Once for each element</button>
<button class="opt" data-correct="false">Once, for the first item</button>
<button class="opt" data-correct="false">Once for each, then drops</button>
</div>
<div class="explain hidden">Adapters are lazy: <code>map</code> only builds a <code>Map</code> struct describing the work. Nothing iterates until a consumer calls <code>next</code>, so the closure never runs and you get <code>warning: unused `Map` that must be used — iterators are lazy and do nothing unless consumed</code>. A chain that does not end in a consumer has not run.</div>
</div>
<div class="q" data-type="recall" data-topic="Iterators">
<p class="topic">Iterators</p>
<p class="prompt"><code>complete</code> needs the task itself; <code>remove</code> needs its index. Which adapter does each want, and which iterator does each start from?</p>
<button class="reveal-btn">Show answer</button>
<div class="answer hidden"><code>complete</code> wants <code>self.tasks.iter_mut().find(|t| t.id == id)</code> — <code>find</code> returns the element, and <code>iter_mut</code> because it assigns to <code>status</code>. <code>remove</code> wants <code>self.tasks.iter().position(|t| t.id == id)</code> — <code>position</code> returns <code>Option&lt;usize&gt;</code>, which is what <code>Vec::remove</code> takes, and plain <code>iter</code> is enough because it only looks. Both then take <code>.ok_or(TaskError::NotFound(id))?</code>.</div>
<div class="grade hidden">
<button data-grade="hit">Got it</button>
<button data-grade="miss">Missed it</button>
</div>
</div>
<div class="q" data-type="mcq" data-topic="Collections">
<p class="topic">Collections</p>
<p class="prompt">Four lines are parsed, the third is corrupt, and you <code>collect::&lt;Result&lt;Vec&lt;Task&gt;, TaskError&gt;&gt;()</code>. What comes back?</p>
<div class="options">
<button class="opt" data-correct="true">One <code>Err</code>, holding that line</button>
<button class="opt" data-correct="false">One <code>Ok</code>, holding three tasks</button>
<button class="opt" data-correct="false">One <code>Ok</code>, holding four results</button>
<button class="opt" data-correct="false">One <code>Err</code>, holding four errors</button>
</div>
<div class="explain hidden"><code>Result</code> implements <code>FromIterator</code>, so collecting an iterator of <code>Result</code>s into a <code>Result&lt;Vec&lt;_&gt;, E&gt;</code> short-circuits: the first <code>Err</code> stops the iteration and becomes the whole answer, and the successful items are dropped. That is the behaviour a save file wants — no half-loaded store. Collect into <code>Vec&lt;Result&lt;Task, TaskError&gt;&gt;</code> instead and you keep all four outcomes.</div>
</div>
<div class="q" data-type="recall" data-topic="Collections">
<p class="topic">Collections</p>
<p class="prompt">Why must a <code>HashMap</code> key implement both <code>Eq</code> and <code>Hash</code>, and why is <code>PartialEq</code> not enough?</p>
<button class="reveal-btn">Show answer</button>
<div class="answer hidden">Lookup is two steps: hash the key to choose a bucket (<code>Hash</code>), then compare for equality inside it (<code>Eq</code>). <code>Eq</code> is a marker trait with no methods that adds one promise <code>PartialEq</code> does not make — every value equals itself. <code>f64</code> breaks that promise, since <code>NAN != NAN</code>, so it implements <code>PartialEq</code> only and cannot be a key. A fieldless enum can promise it, so <code>#[derive(PartialEq, Eq, Hash)]</code> is honest.</div>
<div class="grade hidden">
<button data-grade="hit">Got it</button>
<button data-grade="miss">Missed it</button>
</div>
</div>
<div class="q" data-type="recall" data-topic="Ownership">
<p class="topic">Ownership</p>
<p class="prompt"><code>|task| task.priority</code> in a closure over <code>&amp;Task</code> gives <code>error[E0507]: cannot move out of ... behind a shared reference</code>. What is the cause, and which of the three fixes wins?</p>
<button class="reveal-btn">Show answer</button>
<div class="answer hidden">A map key must be <em>owned</em>, because the map keeps it — but the closure only has a borrow of the task, so reading the field out of it is a move from something you do not own. Fixes: <code>.clone()</code>, <code>#[derive(Clone, Copy)]</code>, or key by <code>label()</code> to get a <code>&amp;str</code>. Derive <code>Copy</code>: a fieldless enum is one small integer, so copying is free, and std does the same for its own small enums such as <code>ErrorKind</code>.</div>
<div class="grade hidden">
<button data-grade="hit">Got it</button>
<button data-grade="miss">Missed it</button>
</div>
</div>
<div class="q" data-type="recall" data-topic="Traits">
<p class="topic">Traits</p>
<p class="prompt">In <code>fn tally&lt;T, K, F&gt;(items: &amp;[T], key: F)</code> with <code>K: Eq + Hash, F: Fn(&amp;T) -&gt; K</code> — why does <code>F</code> have to be a type parameter at all, and what does the <code>where</code> clause buy you?</p>
<button class="reveal-btn">Show answer</button>
<div class="answer hidden">Every closure has its own unique anonymous type, so there is no concrete type to write down; a parameter bounded by the <code>Fn</code> trait is the only way to accept one. The bounds are a two-way contract: callers must supply types that satisfy them, and the body may use only the operations they guarantee — so <code>item.to_string()</code> would not compile without <code>T: Display</code>. That is checked once against the definition, and monomorphisation then compiles one specialised copy per set of types, so the generality costs nothing at runtime.</div>
<div class="grade hidden">
<button data-grade="hit">Got it</button>
<button data-grade="miss">Missed it</button>
</div>
</div>
<div id="summary"><div id="summary-body"></div><p id="summary-total"></p>
<button id="report-btn">Copy report</button><pre id="report-output" class="hidden"></pre></div>
<h2>The drill — 35 minutes, your own crate</h2>
<p>Type it, do not paste it. The bookshelf above is a different program. Keep the
<a href="../reference/rust-syntax.html#iterators">iterators reference</a> open — looking syntax up is free.</p>
<pre><code>cd ~/learn-rust/tasks
cp ../lessons/0008-collections-spec.rs tests/collections.rs
cargo test # 14 new tests fail to compile — that is the starting line</code></pre>
<p>Do not edit anything in <code>tests/</code>. All 32 existing tests must still pass. Target at the end:
<strong>46 passing</strong>.</p>
<h3>Step 0 — the second <code>From</code>, two minutes</h3>
<p>Write the impl that 0007 asked for and then delete the workaround, so <code>?</code> handles io errors
everywhere from here on:</p>
<pre><code>impl From&lt;io::Error&gt; for TaskError { .. } // in error.rs, beside the other
fs::write(path, contents)?; // in save — map_err goes away</code></pre>
<p><strong>Check:</strong> <code>grep -c "impl From&lt;io::Error&gt;" src/error.rs</code> prints <code>1</code>,
<code>grep -c "map_err(TaskError::Io)" src/store.rs</code> prints <code>0</code>, and
<code>cargo test --test persist</code> still passes 8.</p>
<h3>Step 1 — a new module and one generic function</h3>
<p>Create <code>src/stats.rs</code>, declare it in <code>lib.rs</code>, and write <code>tally</code> from
Part 6 — from the signature, not by copying the body. It is nine lines.</p>
<p><strong>Check:</strong> <code>cargo test --test collections tally</code> → 3 passed.</p>
<details>
<summary>Forgotten how a module is declared?</summary>
<p><code>pub mod stats;</code> in <code>src/lib.rs</code>, alphabetically beside the others. Without that line
the file is not compiled at all and you get <code>error[E0432]: unresolved import</code> from the test file —
the same error 0006 showed you.</p>
</details>
<h3>Step 2 — <code>Priority</code> as a key</h3>
<p>Add <code>count_by_priority(&amp;self) -&gt; HashMap&lt;Priority, usize&gt;</code> to <code>Store</code>, as one
line delegating to <code>tally</code>. Let it fail first, read all three errors, and fix them with the derives
they ask for. Seeing E0277 twice and E0507 once, in that order, is the point of the step.</p>
<p><strong>Check:</strong> <code>cargo test --test collections count_by</code> → 3 passed.</p>
<h3>Step 3 — two more methods on <code>Store</code></h3>
<pre><code>pub fn titles_with(&amp;self, priority: Priority) -&gt; Vec&lt;&amp;str&gt;
pub fn remove_completed(&amp;mut self) -&gt; usize</code></pre>
<p><code>titles_with</code> answers “what am I meant to be doing at this priority?”. Hand it a priority and it
gives back the title of every task that carries that priority, in the order the tasks were added. Nothing
matches, and you get an empty <code>Vec</code> rather than an error — an empty answer is a legitimate answer
here. Note the return type: <code>Vec&lt;&amp;str&gt;</code>, not <code>Vec&lt;String&gt;</code>. It hands back
borrows of titles the store still owns, so nothing is cloned, and <code>task.title.as_str()</code> is the
conversion you need.</p>
<p><code>remove_completed</code> is the tidy-up: it deletes every task whose status is <code>Done</code> and
returns how many it deleted. Three details the tests hold you to. The tasks that survive keep their own ids —
you are removing rows, not renumbering them. The id counter is untouched, so the next <code>add</code> carries
on from where it had got to rather than reusing a freed number. And removing nothing is a normal outcome that
returns <code>0</code>, not an error. One <code>Vec</code> method from Part 3 does the removal; the count is
the length before minus the length after.</p>
<p><strong>Check:</strong> <code>cargo test --test collections</code> → 11 of 14 passed.</p>
<h3>Step 4 — rewrite <code>store.rs</code> with what you learned</h3>
<p>Four functions, all currently loops, all one expression each: <code>complete</code> with
<code>iter_mut().find()</code>, <code>remove</code> with <code>iter().position()</code>, <code>save</code>
with <code>map(..).collect::&lt;String&gt;()</code>, and <code>load</code> with
<code>lines().map(str::parse).collect::&lt;Result&lt;Vec&lt;Task&gt;, TaskError&gt;&gt;()?</code>. In <code>load</code>
the <code>if !items.is_empty()</code> guard goes away with the splitter, and the trailing
<code>mut contents = String::new()</code> dance collapses into the <code>match</code> from 0007 returning
a value.</p>
<p><strong>Check:</strong> <code>grep -c "for " src/store.rs</code> prints <code>0</code>,
<code>grep -c "lines()" src/store.rs</code> prints <code>1</code>, and <code>cargo test</code> → 17 + 7 + 8 +
14 = <strong>46 passed</strong>.</p>
<details>
<summary>Stuck on <code>save</code> building a <code>String</code> from an iterator?</summary>
<p><code>String</code> implements <code>FromIterator&lt;String&gt;</code>, so a chain of owned lines collects
straight into one: <code>self.tasks.iter().map(|t| format!("{}\n", t.to_line())).collect()</code>. Annotate the
target — <code>let contents: String = ..</code> — or E0283 will ask you which of several possible types you
meant.</p>
</details>
<h3>Step 5 — two new commands, so the CLI shows it</h3>
<p>Add <code>Stats</code> and <code>Clear</code> to the <code>Command</code> enum and to
<code>Command::parse</code> (the words are <code>stats</code> and <code>clear</code>). The
<code>match</code> in <code>main</code> will refuse to compile until both are handled — that is
<code>E0004</code>, the same non-exhaustive-match error from 0006, doing its job again.</p>
<p><code>stats</code> must print the three priorities in a fixed order, because hash map iteration order is
arbitrary. Loop over <code>[Priority::High, Priority::Medium, Priority::Low]</code> and ask the map for each,
with <code>counts.get(&amp;p).copied().unwrap_or(0)</code> so an absent priority prints <code>0</code> rather
than vanishing.</p>
<p><strong>Check — a real session, run today against the reference implementation:</strong></p>
<pre><code>$ cd $(mktemp -d)
$ run add "buy milk" high
added task 1
$ run add "call bank"
added task 2
$ run add "water plants" low
added task 3
$ run done 1
completed 1
$ run stats
high 1
medium 1
low 1
$ run clear
cleared 1 completed
$ run list
2 [todo] call bank (medium)
3 [todo] water plants (low)
$ cat t.txt
2|todo|medium|call bank
3|todo|low|water plants
$ run stats ; echo $?
high 0
medium 1
low 1
0</code></pre>
<p>(<code>run</code> above is
<code>TASKS_FILE=t.txt cargo run -q --manifest-path ~/learn-rust/tasks/Cargo.toml --</code>.) That last
<code>high 0</code> is <code>.copied().unwrap_or(0)</code> earning its place: the completed high-priority task
is gone, so the map has no <code>High</code> entry at all, and the absence prints as a zero instead of a
missing line.</p>
<h3>Then stop</h3>
<p>Not today: <code>fold</code> and <code>zip</code>, <code>BTreeMap</code> (sorted keys — the right answer if
you ever want <code>stats</code> ordered without hard-coding), <code>impl Iterator for</code> your own type, and
<code>itertools</code>. Each is a small step from here, and none of them is on the path to the next gap.</p>
<h2>What this closed</h2>
<p>Chapters 8 and 13 move to <em>produced</em> on the <a href="../reference/book-coverage.html">coverage
map</a>, and 10.1 opens with a real generic function of your own rather than a book example. What is left before
the job-ready floor is short:</p>
<ol>
<li><strong>ch 11 — writing your own tests</strong> (lesson 0009). You have now consumed 46 of my tests and
written zero. Test-writing is a first-round interview question, and it is the last big gap in the book's core.</li>
<li><strong>ch 10.3 — lifetimes</strong>, as reading practice. You wrote one today without noticing:
<code>titles_with</code> returns <code>Vec&lt;&amp;str&gt;</code> borrowed from <code>&amp;self</code>, and
elision filled in the annotation for you.</li>
<li>Then <code>serde</code> → <code>axum</code>, where the trait work from 0005–0008 starts paying rent.</li>
</ol>
<h2>Take it outside</h2>
<p>Here is a question with genuine disagreement behind it, which makes it a good one to ask people rather than
docs. Your <code>tally</code> takes <code>&amp;[T]</code>. Most experienced Rust developers would write it to
take <code>impl IntoIterator&lt;Item = T&gt;</code> instead, so it accepts a <code>Vec</code>, an array, a
<code>HashSet</code>, or any chain of adapters — not only a slice. Post <code>tally</code> on
<a href="https://users.rust-lang.org">users.rust-lang.org</a> (Code Review category) and ask whether the
<code>IntoIterator</code> version is worth the extra signature complexity for a small crate, and where they
personally draw that line. The answers will teach you more about idiomatic API design than any chapter, because
it is a taste question and the book cannot have taste for you.</p>
<h2>The five sentences worth keeping</h2>
<ol>
<li>Adapters are lazy and return iterators; consumers do the work. A chain that does not end in a consumer never
ran.</li>
<li><code>iter</code> borrows, <code>iter_mut</code> borrows mutably, <code>into_iter</code> takes ownership —
the <code>&amp;self</code>/<code>&amp;mut self</code>/<code>self</code> choice, one element at a time.</li>
<li><code>collect</code> builds any <code>FromIterator</code> type, so an iterator of <code>Result</code>s
collects into one <code>Result&lt;Vec&lt;_&gt;, E&gt;</code> that short-circuits on the first error.</li>
<li>A <code>HashMap</code> key needs <code>Eq + Hash</code>; <code>*map.entry(k).or_insert(0) += 1</code> is the
counting idiom; iteration order is arbitrary, so impose your own before printing.</li>
<li>A generic function's <code>where</code> clause is a contract checked once against the definition, and
monomorphisation means the generality is free at runtime.</li>
</ol>
<footer>
<p><strong>Primary source:</strong> The Rust Book
<a href="https://doc.rust-lang.org/stable/book/ch13-02-iterators.html">13.2 — Processing a Series of Items with
Iterators</a>, then <a href="https://doc.rust-lang.org/stable/book/ch08-03-hash-maps.html">8.3 — Hash Maps</a>.
After the drill, skim the method list on
<a href="https://doc.rust-lang.org/std/iter/trait.Iterator.html">std::iter::Iterator</a> — not to memorise it, but
so you know what is there to look for. It is the single highest-value page in std.</p>
<p>Previous: <a href="0007-files-and-fromstr.html">0007 — Files, io::Error, FromStr</a> ·
<a href="0006-your-own-error-type.html">0006 — Your own error type</a> ·
<a href="0005-traits-display-and-errors.html">0005 — Traits, Display, errors</a><br />
Reference: <a href="../reference/rust-syntax.html#iterators">Iterators</a> ·
<a href="../reference/rust-syntax.html#collections">Collections</a> ·
<a href="../reference/rust-syntax.html#traits">Traits &amp; generics</a> ·
<a href="../reference/book-coverage.html">Coverage map</a></p>
<p><strong>Ask me things.</strong> Bring the compiler output verbatim — step 2 is meant to fail three times, and
step 4 is the first time you will rewrite working code purely for shape, which is a different kind of
uncomfortable. If a paragraph did not land, name it; that is my fault to fix, and cheaper to fix now than
mid-drill.</p>
</footer>
<script src="../assets/quiz.js"></script>
</body>
</html>