Skill: write-like-james → write-mcqs, review-questions (rules read; nothing authored) @ 84b31cb
Grade 3 Science end-of-strand tests: a format proposal for external validation
From James Moore's Grade 3 Science build, 30 September 2026 (revised the same evening against Becky's "A - Mastery Gates" and "X - Blueprint Creation and Form Review"). For Becky (Alpha assessment) and James. Nothing in this document is authored test content; it is the format we propose to build, set beside David's Alpha Standardized Science Grade 3 blueprint and Becky's design rules, with the questions we need answered before any item is written.
What the course is. Grade 3 Science is one self-paced course on TimeBack Scroll: 183 one-idea lessons in 42 small topics, in four strands taught in order (Physics, Matter, Earth & Space, Life). Every topic ends with a PP100 drawn from a bank of about fifty plain items, one atom (knowledge component) per item. The course follows TEKS elementary science, NGSS 3-5, Florida NGSSS and Arizona, so it teaches content the NGSS Grade 3 band leaves to other grades: matter and its states, energy, light, sound, the Sun, stars and solar system, soil, natural resources, food chains.
What the tests are for. Four end-of-strand tests, one after each strand, decide whether the child is ready to move on to the next strand. In Becky's terms they are within-course instructional gates, not endpoint mastery assessments. The Grade 3 endpoint already exists as David's family of ten live forms (G3.22 to G3.31); we do not build one.
What we ask Becky to validate, and by when. Four things, in this order: (1) the assessment claim and the classification of the strand test as a within-course gate (section "The gate in Becky's terms"); (2) the synthetic blueprint written in her ten-step shape (section "The blueprint"), in particular the pitch, the strata and the retake policy; (3) the eight questions below, which her documents raise or leave open; (4) once she has signed off the format, five form plans for the Physics gate (slots only, no items) and then one authored form, each reviewed against Steps 8 and 9 of her process before the other forms are written. We propose the format and claim this week and the Physics form plans within a week of her sign-off; the dates are hers and James's to set. Nothing goes live on her side of the gate.
Questions
1. Is 30 items the right length for a fixed-form strand gate, with the cut of 27 chosen alongside it?
David's number; three-quarters of the 36-item Grade 3 endpoint form. Becky's design rule 4 says length and cut are chosen together for a fixed form; her Economics example is 30 items, 27 to pass, about 30 minutes, untimed.
Recommendation: 30 items, untimed, 27 to pass, a ready child finishing in about 30 minutes. Not an hour. Choices: 30 / 24 (shorter, same cut ratio) / other.
2. Should Physics, with 64 atoms in 15 topics, be one gate or two?
Becky's rule 2 sizes a within-course gate at about a dozen knowledge components; her Issue 1 says a hundred is an exam, not a gate. The topic PP100s already gate 3 to 8 atoms each; the strand test samples across 22 (Matter) to 64 (Physics) atoms.
Recommendation: one gate per strand, sampling by topic strata, with Physics split into two gates (Motion and Forces, Topics 1 to 10; Sound, Energy and Light, Topics 11 to 15) if Becky judges 64 too broad to gate at 30 items. Choices: one Physics gate / two.
3. Besides 27 of 30, does the gate require at least one correct item in every topic stratum?
Becky's rule 4: if mastery requires several essential facets, set a requirement for each rather than averaging them away. A child could score 27 by missing all three items of one topic.
Recommendation: yes: 27 of 30 and no topic stratum at zero correct; a stratum at zero routes the child to that topic's lessons and PP100, whatever the total. Choices: yes / total only / a per-stratum floor only for prerequisite-heavy topics.
4. Is the retake policy right: at most two retests, each on an unseen fixed form after the missed topics' instruction, and what happens after a third fail?
Becky's Principle 9: unlimited retakes destroy the cut; failure must change the route; the pass rule is validated across the whole retake policy. Principle 6: a retest shows only unseen items.
Recommendation: three attempts in all (Forms A, B, C); after a third fail the child redoes the failed topics' lessons and PP100s in full and a teacher decides. Choices: as recommended / two attempts / other.
5. Which fail scores route the child straight back to instruction, and which send them through the topic PP100s first?
Becky's rule 5 asks for both bands to be named before launch.
Recommendation: below 21 of 30 (70 per cent): back to the lessons of every topic with a missed item; 21 to 26: redo the PP100 of each topic with a missed item, then the next form. Choices: as recommended / other bands.
6. May one-sentence constructed items sit in the gate before Scroll's marking has been checked against human marking?
Becky's Principle 8 says free response is often needed because explaining is part of the claim, with a binary mark and an excessively clear stem; marker reliability is a real error source. Scroll marks free response with an AI grader we have not yet tested on a child's sentence.
Recommendation: two one-sentence items per form, binary-marked, but only after a sample of 50 child answers has been marked by the grader and by James or Olivia and the agreement recorded; until then those two slots are MCQs. Choices: as recommended / never in the gate / in from the start.
7. Which Alpha item types can TimeBack Scroll render, and are they familiar enough from practice not to count as unfamiliar controls?
Scroll's contract on our disk lists MCQ, AI-graded free response, match pairs, cloze, sequence and sort; no multi-select, no drag-to-build graph, no hot-text, no numeric-entry type. Becky's Principle 5 names unfamiliar controls as construct-irrelevant difficulty.
Recommendation: use only Scroll types the child has met in the practice form; map the rest as in Card 3; ask the Scroll team today. Choices: as mapped / wait for Scroll's answer before fixing the mix.
8. Who certifies the Grade 3 content the AlphaTest Grade 3 family leaves out?
David's spec tests only what the old Alpha Grade 3 course taught and excludes matter, energy, light, astronomy, soil and food chains. James's course teaches all of them at Grade 3 (about 80 of 183 atoms).
Recommendation: the strand gates certify readiness on that content now; David re-derives the Grade 3 family's testable-content list from James's course when it is live, under his own Q8/Q9 rule. Choices: as recommended / strand gates only / revise the family first.
Settled by Becky's documents, so not asked (answer and source):
- The pitch: a gate is set at readiness to proceed, never at the course's ceiling; items are ones a ready child answers reliably; nothing above the mastery level is served (A, Principles 2 and 4, Issue 3). So the strand test does not copy David's 40 per cent DOK 3 posture, which belongs to the endpoint's rigour-superset claim.
- The bar: 90 per cent is a convention, not a measurement argument; whatever the cut, items must be pitched to match it, and a lower cut is never a remedy for flawed items (A, Principle 8). We keep 27 of 30 and pitch to it.
- Protected items: every gate item is unseen, unhinted, single-attempt, from a pool never used for practice (A, Principle 6). So the practice form draws from a separate pool, not the gate bank.
- Marks: binary, one mark per item; a two-part pair is two items (A, Issue 12; rule 3).
- Cumulative sampling: a gate's claim is one thing (rule 1); the strand gate's claim is one strand, so the Life gate does not sample Physics.
- Tags: easy / medium / hard describe complexity, not difficulty; no invented numerical difficulty targets (A, Issue 3; X, Step 5).
- Assembly: no knowledge component is fixed to a slot across forms; sample within strata; build the form plans together (X, Step 3 and Step 7).
- Validation: one form, then at least five varied forms or form plans, compared for coverage, demand and meaning (X, Principle 6, Steps 8 and 9); live monitoring indicators named before launch (X, Step 10).
Applied without asking (rulings already given): static Form A (James, today); the 90 per cent bar; Matter gets a gate like the other three; tests carry no hints or frames; every key passes blind review; an end-of-strand test gets the examiner pass against released forms before its blind read; David's untimed rule (Q7) and his independent scoring of two-part items (Q6).
Becky's WorkFlowy guidance: reconciliation
Read in full: reference/becky_WF_A_Mastery_Gates_2026-09-30.md and reference/becky_WF_X_Blueprint_Creation_2026-09-30.md, including the Economics, New York, MAP and Grade 5 Mathematics examples.
Where the proposal already followed her logic: protected, blind-reviewed items; one atom per item; fresh instances; a fixed form with the cut chosen with its length; a strand-scoped claim; a synthetic blueprint informed by external forms; released items as anchors, never copied; live-data review after launch.
Where it conflicted, and what we change:
- Pitch. We had adopted David's endpoint DOK mix (about 40 per cent DOK 3). Her Principles 2 and 4 and Issue 3 place a within-course gate at readiness, with items a ready child answers reliably. Changed: DOK 2 is the mastery level; DOK 3 only where the standard's own practice is DOK 3 (planning a fair test, arguing from evidence); at most 3 DOK 1.
- Retakes. We had "no cap on attempts; from the third attempt items may recur". Her Principles 6 and 9 forbid both. Changed: three attempts in all, every one on an unseen form, instruction between them, fail bands named.
- The practice form. We had it drawn from the gate bank. Her Principle 6 keeps practice and gate pools apart. Changed: a separate practice pool; the practice form is also where every item type is first met, so no control is unfamiliar at the gate (Principle 5).
- Free response. We had it off the scored form. Her Principle 8 wants produced answers where explaining is part of the claim, binary-marked, with excessively clear stems. Changed: two one-sentence items per form, once the marker has been checked (question 6).
- Per-facet requirement. We had a total only. Her rule 4 asks for a requirement per essential facet. Proposed: a per-stratum floor (question 3).
- The "format on-ramp". David's phrase makes format preparation a secondary goal of the gate. Her Principle 10 allows a secondary goal to shape delivery but never what counts as passing, and says external-test practice belongs in a separate practice instrument. Changed: the practice form carries the on-ramp; the gate's formats are chosen for the evidence they give, not for practice.
- Forms and bank. We had four fixed forms (five for Physics) and coverage of every atom twice across them. Her Step 7 sizes fixed forms to the known number of attempts and forbids fixing a component to a slot. Changed: three protected fixed forms per gate, planned together, plus a separate practice pool of about 60.
- Words. Easy / medium / hard are complexity tags from here on.
The proposed format beside David's Grade 3 spec
Cells marked per Becky's A or per Becky's X changed on her documents.
| David's Grade 3 endpoint form (G3.22 to G3.31) | Proposed end-of-strand gate (Grade 3) | |
|---|---|---|
| Kind of instrument | course/grade endpoint mastery assessment | within-course instructional gate on one strand (per Becky's A, Purpose and Principle 10) |
| Items per form | 36 nominal (34 to 38) | 30 |
| Marks | about 41 to 44 (multi-select and two-part 2 pt) | 30, one binary mark per item; a two-part pair is two items (per Becky's A, Issue 12) |
| Time | untimed; about 35 to 50 minutes expected | untimed; about 30 minutes for a ready child |
| Scope | all 18 NGSS Grade 3 PEs at least once, plus taught extensions | one strand, sampled by topic strata in proportion to atoms; no atom compulsory or fixed to a slot (per Becky's X, Step 3) |
| Cumulative sampling | none at Grade 3 (25 to 30 per cent at Grade 5) | none |
| Pitch and DOK mix | 0 to 2 DOK 1 / 19 to 20 DOK 2 / 14 to 15 DOK 3 (rigour superset) | 0 to 3 DOK 1 / 20 to 22 DOK 2 / 6 to 8 DOK 3; DOK 3 only where the standard's practice is DOK 3; at least one DOK 2 in every stratum (per Becky's A, Principles 2 and 4, Issue 3) |
| Clusters | 7 phenomenon stimuli × 2 to 3 items, about 50 per cent | 5 to 6 stimuli × 2 to 3 items, about 40 to 50 per cent; a stimulus is one short case with one figure |
| Reading load | stimulus ≤ 100 to 150 words; sentences ≤ 15 words | stimulus ≤ 80 words, one figure; sentences ≤ 12 words; one judgement per item (per Becky's A, Principle 5, young children) |
| Item types | 15 MC · 5 multi-select · 3 two-part pairs (6) · 3 gap-match · 2 inline choice · 2 hot-text · 1 match or order · 2 numeric entry | 14 MC · 3 two-part pairs (6) · 3 match or sort · 2 cloze · 1 sequence · 2 numeric entry · 2 one-sentence constructed (Card 3; per Becky's A, Principle 8 and rule 3 for the constructed items) |
| Free response | none (Q12) | 2 one-sentence items, binary-marked, excessively clear stems; MCQs in those slots until the marker check passes (per Becky's A, Principle 8) |
| Vocabulary | 6 to 8 vocabulary-bearing DOK 2 items | 4 to 6, the strand's coined terms inside application items |
| Complexity tags | easy / medium / hard | complexity, not difficulty; no numerical difficulty targets before live data (per Becky's A, Issue 3) |
| Mastery bar | 89.5 per cent of points, one attempt per form | 27 of 30, and (proposed) no topic stratum at zero correct (per Becky's A, rule 4; question 3) |
| Retakes | retake on another form; ten forms for volume | at most two retests, each on an unseen form after the missed topics' instruction; fail bands named (per Becky's A, Principles 6 and 9, rule 5) |
| Forms and bank | 10 live forms per grade, no stimulus or template reused | 3 protected fixed forms per gate (A, B, C), planned together; a separate practice pool of about 60 (per Becky's A, Principle 6; X, Step 7) |
| Sourcing | released items first, generation last | the same order, filtered to the strand's atoms; our own items where none exist (most of Matter, energy, light, sound) |
| Read-aloud | yes at Grade 3 (stems, prompts, options) | yes if Scroll offers it; note Becky's caution that audio swaps a reading load for a listening-and-memory load |
| Validation | plan audit → authoring → blind solve → cross review → validator → human walk | the same stages in our tools; one form then five varied form plans compared; Becky's gate before live; monitoring indicators named before launch (per Becky's X, Principle 6, Steps 8 to 10) |
The gate in Becky's terms
The call: the strand test is a within-course instructional gate on a bundle of 22 to 64 knowledge components; that fixes its pitch, its threshold logic, its stopping rule and its failure routing.
- Claim (rule 1). For a Grade 3 child who has completed a strand, a pass supports the conclusion that they are ready to begin the next strand without being held back by missing knowledge of this one. James and the child's guide use it to route the child on or back. A false hold costs the child restudy of what they know; a false promotion costs little in Grade 3 Science because strands are only loosely dependent (Life leans on the Sun and on Earth's pull as working knowledge) and the topic PP100s remain in place. So the gate can be short and sample.
- Pitch (Principles 2, 4, 5). Items a ready child answers reliably: fresh cases of the strand's ideas at the level the lessons taught, DOK 2 as the norm; DOK 3 only where the atom's behaviour is itself a DOK 3 practice; never the hardest on-grade question we can devise. No long stimuli, no bundled judgements, no unfamiliar controls.
- Threshold (Principle 8, rule 4). 27 of 30 is kept because James's course-wide convention is 90 per cent and the pitch is chosen to match it; the per-stratum floor (question 3) stops a whole topic being averaged away.
- Stopping rule (rule 4). A fixed form of 30; every path ends: pass, or a third fail and a teacher's decision. Sequential stopping is a later efficiency once the bank carries per-atom tags and live data.
- Failure routing (Principle 9, rule 5). Below 21: back to the lessons of every missed topic. 21 to 26: the PP100 of every missed topic, then the next unseen form. A stratum at zero: that topic's lessons, whatever the total. A waiting period alone never counts as remediation.
- Secondary goals (Principle 10, rule 7). Format familiarisation for the endpoint is a secondary goal; it shapes which item types appear, never what counts as passing, and the practice instrument carries the practice.
- Owner, version, review (rule 6). Owner James Moore; version 0.1 (this proposal); consequence level low (routing within a self-paced course); review triggers: any change to the strand's atoms, the platform's item types, the cut, or live data showing a false-hold or false-pass pattern.
What changes: the gate is written up as above in the strand test's record; the PP100 (about a dozen atoms per topic cluster) stays the fine-grained gate that certifies each atom.
The blueprint (in the shape of Becky's "X - Blueprint Creation", ten steps)
A synthetic blueprint: no single external assessment is the target. David's Grade 3 spec and the released Grade 3 forms (Tennessee TCAP, Louisiana LEAP, Arkansas ATLAS practice) inform the pitch and the item grammar; James's course defines the content.
Step 1: the assessment claim
For Grade 3 Science students who have completed a strand, the gate result is intended to support the conclusion that they can apply that strand's taught ideas to fresh cases at the level the lessons taught, for the purpose of deciding whether they move on to the next strand or return to named topics, in relation to James Moore's Grade 3 Science course (TEKS, NGSS 3-5, Florida, Arizona). It is not intended to support conclusions about mastery of every atom (the topic PP100s do that), about readiness for the Grade 3 endpoint AlphaTest, about performance on any state test, or about hands-on investigation skill.
Step 2: external assessments and how they inform the blueprint
- David's Alpha Standardized Science Grade 3 spec v1.1 (the endpoint): source of the item grammar (phenomenon clusters, two-part evidence, matching, numeric entry), the untimed rule and the 89.5 per cent bar. Not the source of the pitch: its 40 per cent DOK 3 posture serves an endpoint rigour claim, not readiness.
- Tennessee TCAP Grade 3 (25 to 26 items, 50 minutes, MC plus multi-select): a floor for length and a model for data-bearing items.
- Louisiana LEAP Grade 3 (36 items, six four-item sets): the model for cluster shape and two-part evidence.
- Arkansas ATLAS Grade 3 (six clusters plus ten standalones, untimed): the model for an untimed, cluster-based Grade 3 sitting.
- MAP Growth Science Grade 3 (40 to 43 items, median 32 to 37 minutes): the stamina anchor; not a model for pitch, because MAP pitches at 50 per cent success.
- Released Grade 5 forms (STAAR, Florida SSA, AZSCI): anchors for the TEKS and Florida content no Grade 3 form tests; not on our disk yet.
Similarity recorded qualitatively: closest to LEAP in shape, to TCAP in length, to none in content (the strand scope is ours).
Step 3: content coverage model
- Eligible content: every atom of the strand as taught (the registers are the source of truth); Scientific Reasoning atoms taught inside a strand's topics are eligible there.
- Excluded: nothing within the strand; everything outside it.
- Compulsory: no atom is compulsory or fixed to a slot; every topic stratum appears on every form.
- Strata: the strand's topics, with allocations in proportion to atoms, floor one item per topic. Physics (30 items over 15 topics): Topics 1 and 12 → 4 each; Topic 2 → 3; Topics 3, 5, 7, 9, 11, 13, 14 → 2 each; Topics 4, 6, 8, 10, 15 → 1 each. Matter (6 topics): 5 / 5 / 4 / 4 / 6 / 6. Earth & Space (10 topics): 4 / 4 / 4 / 2 / 3 / 3 / 3 / 2 / 3 / 2. Life (11 topics): 3 / 3 / 2 / 4 / 2 / 3 / 2 / 2 / 3 / 2 / 4. (Two-part pairs count two toward their stratum.)
- Weighting (Becky's Principle 8): topics later grades lean on hardest carry the higher allocations already (Forces, Energy, states of matter, weather, life cycles); if James names a topic as essential, its allocation rises rather than any topic being excluded.
- Coverage control at topic level; atoms sampled at random within the stratum, unseen first.
- Cross-walk for Becky (codes read from the registers): Physics reaches 3-PS2-1 to 3-PS2-4 and TEKS 3.5A, 3.7A, 3.7B, 3.8A, 3.8B; Matter reaches TEKS 3.6A to 3.6D and Florida SC.3.P.8 and 9 (no NGSS Grade 3 PE); Earth & Space reaches 3-ESS2-1, 3-ESS2-2, 3-ESS3-1, 3-LS4-1 and TEKS 3.9 to 3.11; Life reaches 3-LS1-1, 3-LS2-1, 3-LS3-1, 3-LS3-2, 3-LS4-3, 3-LS4-4 and TEKS 3.12, 3.13. Fourteen of David's 18 PEs sit in the course; 3-LS4-2 and 3-5-ETS1-1 to 3 do not.
Step 4: item mix and cognitive demand
- 30 items, one mark each, no penalty for wrong answers.
- Permitted types (Scroll's six): MCQ (four options), two-part evidence pair (two MCQs on one stimulus), match pairs, sort, cloze with a word bank, sequence (life cycle only, never presorted), numeric entry (an exact number from a table or instrument), one-sentence constructed response (binary mark). Multi-select, hot-text and drag-to-build are out until Scroll renders them.
- Type-by-content: numeric entry only in strata with a measured quantity (Topics 2, 17, 18, 22, 23); sequence only for life cycles (Topic 40) and state changes (Topic 21); constructed response only on atoms whose behaviour is "explain".
- DOK controlled across the form: 0 to 3 DOK 1, 20 to 22 DOK 2, 6 to 8 DOK 3; at least one DOK 2 in every stratum; DOK 3 only where the PE or TEKS verb is plan, argue, explain with evidence, or compare two data sources.
- Fit: 30 items at about a minute each plus two short sentences is about 30 to 35 minutes for a ready child.
Step 5: acceptable item boundaries (task sketches, not items)
Bands as in Becky's examples: too low / on target lower boundary (DOK 1) / on target secure (DOK 2) / on target upper boundary (DOK 3) / too high. One stratum per strand here; the rest follow after her sign-off of the format, before any form plan.
- Physics, Topic 9 Balanced forces. Too low: "What do we call forces that cancel out?" (the word is in the lesson title). Lower boundary: two children push a box from opposite sides with the same strength; state what the box does. Secure: a drawing of a hanging lamp with two arrows of equal length; pick the reason it hangs still. Upper boundary: a table of three tug-of-war rounds with each team's pull; pick the round the rope moved and the evidence that shows it. Too high: add force arrows as numbers and find the net force (Grade 6 and above).
- Matter, Topic 21 Heating and cooling. Too low: "Ice is a solid. What is ice?" Lower boundary: a puddle is gone by afternoon on a sunny day; pick the change's name from the four state-change words. Secure: a thermometer drawing reads 3 degrees on a dish of water outdoors overnight in winter; predict what the water is like in the morning and why. Upper boundary: two identical ice cubes, one on a metal tray, one on a wooden board, with melting times; pick the claim the data supports and the evidence line. Too high: the melting point of any substance other than water, or the particle model of a solution.
- Earth & Space, Topic 22 Measuring the weather. Too low: "Which tool measures temperature?" with a thermometer drawn beside the item. Lower boundary: read a rain gauge drawing to the nearest 2 millimetres (numeric entry). Secure: a week's temperature table; pick the day to put on a coat and the reading that says so. Upper boundary: two weeks' records from two towns; pick which town had the wetter week and the two readings that decide it. Too high: forecasting from air pressure or the water cycle (Grade 4).
- Life, Topic 40 Life cycles and family likeness. Too low: "Do puppies look like their parents?" Lower boundary: order four drawn stages of a frog's life (sequence, unsorted). Secure: a litter of kittens with one white kitten and two grey parents; pick the statement about offspring that fits. Upper boundary: a table of two bean plants grown from seeds of one parent in different light; pick which differences came from the parents and which from where they grew, with the evidence. Too high: genes, dominant and recessive traits, or natural selection over generations.
Step 6: whole-assessment requirements
- Delivered on TimeBack Scroll in one sitting; fixed order; no revisiting an answered item; an interrupted sitting resumes where it stopped.
- 30 items, 30 marks, binary; pass at 27 and no stratum at zero (question 3).
- Untimed; designed for about 30 minutes.
- Closed book; no notes; no calculator needed (numbers stay small); read-aloud if the platform offers it, recorded as an accommodation.
- Every item carries one atom, one topic stratum, one complexity tag (DOK), one type, and its released analogue where one exists.
- Clusters: 5 to 6 stimuli per form; one short case, one figure, no more than three items; an item never reveals another's answer.
- No hints, no frames, no reference table beside the item it answers; feedback shown only after the sitting closes.
- Retakes: at most two, each on an unseen form, only after the routing in Step 6's bands; three attempts in all.
Step 7: production and assembly
- Fixed-form route for the gate (a known, limited set of forms): three form plans per gate (A, B, C) created together so atoms are spread within strata and no atom is fixed to a slot; each slot names stratum, atom, DOK, type and stimulus group; plans reviewed against this blueprint before any item is authored; an item authored per slot through write-mcqs (fresh instance, design-move inventory, one atom), then blind read, review-questions, examiner pass, sweeps.
- Item-bank route for the practice form: a separate pool of about 60 per gate, the same strata and types, with feedback lines; drawn 30 at a time; never used in the gate.
- Bank depth is a per-stratum property; a stale (already seen) item served at the gate is a fault to report, never a design choice.
Step 8 and Step 9: one form, then several
- Author and review one complete form (Physics A) through compliance, representation and practicality before the others.
- Before authoring, generate five deliberately varied form plans from the blueprint and compare them for stratum coverage, DOK, type balance, expected challenge and meaning; if two plans differ materially, the blueprint is tightened first.
Step 10: monitoring from the first live sitting
- Per item: percentage correct among course completers (target: a ready child gets it right reliably; any item below about 75 per cent among children who passed their topic PP100s is reviewed), omissions, long response times, distractor use.
- Per form: pass rate, score distribution near 27, stratum-zero holds, retake counts and outcomes, agreement of the AI marker with human marking on the two constructed items.
- Across measures: relation to the child's topic PP100 results and to their next strand's PP100s.
- Owner James; reviewer Becky; first review after 50 completers per form; any blueprint revision is versioned and re-validated by Steps 8 and 9.
Card 3: item grammar, and what Scroll must confirm
The call: build the Alpha item grammar on the six types Scroll renders, introduce every type in the practice form first, and ask the Scroll team about the rest today.
| Alpha type (per 36) | On Scroll (per 30) | How it renders | Confirm with Scroll |
|---|---|---|---|
| MC single, 4 options (15) | 14 MC | mcq, four options | none |
| Multi-select, 5 options, 2 pt (5) | 0; the demand moves into two-part pairs and sort items | no multi-select type | a select-all type |
| Two-part evidence pairs (3 pairs) | 3 pairs = 6 items | two mcq rows sharing one stimulus, repeated on each row; scored independently | whether Part B can be locked to Part A |
| Gap-match incl. graph construction (3) | 3 match or sort (match_pairs, sort_categories) | no drag-to-build graph | a gap-match or graph type |
| Inline choice (2) | 2 cloze with a word bank | cloze | none |
| Hot-text (2) | 0; asked as an MC whose options are the sentences | no hot-text type | a hot-text type |
| Match or order (1) | 1 sequence | sequence, never presorted | none |
| Numeric entry (2) | 2 numeric entry as AI-graded frq with an exact accept list | no numeric type | a numeric type; whether the grader honours an exact list |
| Constructed response (0 at Grade 3) | 2 one-sentence items, binary (question 6) | AI-graded frq with one-point marking guidance | marker agreement check first |
Why:
- David's message: "the thing I'd hold onto isn't the length, it's the item grammar."
- Becky's Principle 5: an unfamiliar control is construct-irrelevant difficulty. Every type the gate uses appears in the practice form and in the topics' mixed practice first.
- A phenomenon cluster is two or three rows sharing one stimulus, the stimulus repeated on every row because Scroll serves rows singly.
What changes: the export tool needs a fixed-form gate and a protected gate bank beside the PP100 bank; both are contract questions for the Scroll team, listed in the placeholders file.
What changes for the course
- The PP100 stays the fine-grained gate: about a dozen atoms per topic cluster, one atom per item, mostly plain MCQs with typed production items. It certifies each atom.
- The strand test is a sampling gate on readiness for the next strand, pitched at the taught level, in the Alpha item grammar the child has already met in practice.
- The practice form carries the format on-ramp, from a separate pool, with feedback.
- The AlphaTest stays the endpoint: no format practice is designed for it beyond the above.
- No gate item is authored until Becky's format gate is passed. Today's four Form A agents are stopped; their spec (24 MCQ, 4 typed, 2 free response) is withdrawn.
What sits at the end of each strand to prepare the child
The call: a practice form from its own pool and two free revision pages; nothing else now.
| Resource | What it is | Cost | Recommendation |
|---|---|---|---|
| Practice form | one drawn form in the gate's grammar, from a separate practice pool of about 60, a feedback line under every item | one pool's authoring per gate | build, after the gate; every gate item type is met here first (per Becky's A, Principles 5 and 6) |
| Topic rule sheets | one page per topic: each lesson's rule sentence in teaching order, built from the lesson files | none; the renderer already builds this block | build now |
| Summary-video playlist | the strand's topic summary videos in order | none | build now |
| Key-terms typing set | typed recall over the strand's coined terms | a day's authoring per strand | hold: the PP100 mixed practice already recalls terms |
| Missed-item replay | the child's wrong items served again with feedback | needs answer data from the platform | ask the Scroll team; note it may only use practice-pool items, never gate items (per Becky's A, Principle 6) |
Placeholders
grade3/strand_tests/PLACEHOLDERS.md lists the four gates (Physics, Matter, Earth & Space, Life) with status "format under validation with Becky", their position in the course tree (after Topics 15, 21, 31 and 42), and the intended values of the ledger rows the TimeBack hand-off should carry.
What happens next
James forwards this page to Becky.
Becky and James answer the eight questions; her verdict on the claim, the pitch and the retake policy comes first.
The Scroll team answers Card 3's rendering questions.
Then five Physics form plans (slots only), reviewed against Steps 8 and 9, then one authored form.
David's message to James, and Becky's words, verbatim (30 September 2026)
David: "30 questions per end-of-domain test feels right — about three-quarters of a full end-of-grade form (our G3 forms run 36 items; middle school 40). The thing I'd hold onto isn't the length, it's the item grammar. Build them from the same blueprint DNA as the Alpha tests: phenomenon-driven clusters (the passage-type experience), at least one numeric-entry item, drag-and-drop/matching, two-part evidence items, and a DOK mix weighted toward 2s and 3s; keep the 90% mastery bar so 'passed' means the same thing at every rung. The end-of-unit layer becomes the format on-ramp — leave clear simple MCQs for initial mastery and move the format-preparation job to these tests, and let the AlphaTest be pure validation. The end-of-year Grade 3 test already exists as a family: ten live G3 forms (G3.22–31) built to one spec from released items (STAAR, TN TCAP, AR ATLAS, NY ILS, KY KSA, WI Forward, LA LEAP, GA Milestones, NAEP, TIMSS; NWEA MAP percentile anchor). You can change this spec for the end-of-domain tests or use it as-is."
Becky, relayed: "It is very important you don't just write something and expect to change it in the future. External validation is critical. And once a mastery test is in place it is extremely difficult to change."
James's ask, in chat (30 September, about 15:15): end-of-strand tests for Physics, Earth & Space, Life (static; 90 per cent on ours means 90 per cent on any equivalent state test; a large bank for retests; some text entry and short free response beside MCQ); a plan for length and item count ("an hour?"); a proposal for an end-of-strand practice test and revision resources.
Evidence on disk, and what is not
Becky's pages (exports): reference/becky_WF_A_Mastery_Gates_2026-09-30.md (Principles 1 to 10, design rules 1 to 7, the PP100 revision Issues 1 to 12); reference/becky_WF_X_Blueprint_Creation_2026-09-30.md (Principles 1 to 9, Steps 1 to 10, the Economics, New York, MAP and Grade 5 Mathematics examples).
David's blueprint folder, reference/alpha_test_blueprints/: G3_TEST_DESIGN_SPEC_v1.md (36-item form; 18 PEs; DOK ≤5 / ~55 / ~40; 7 clusters; item-type table; 89.5 per cent bar; Q1 numeric entry only, Q6, Q7 untimed, Q8/Q9 testable content, Q12 no FRQ; 2026-08-18 read-aloud amendment); G3_RESEARCH_DOSSIER_2026-08-07.md (TCAP 25 to 26 items, 50 minutes; LEAP 36 items, two 70-minute sessions; ATLAS 6 clusters + 10, about 75 minutes; MAP 40 to 43 items, median 32 to 37 minutes; Iowa 30 items in 35 minutes; NAEP about 90 seconds an item; Georgia DOK bands); G4_G5_DECISION_CHECKPOINT_2026-08-10.md (Grade 5 cumulative band); G6G8_DECISION_CHECKPOINT_2026-08-11.md (40-item forms; false-fail arithmetic); build_refs_Item_Writing_Rules.md; EXEMPLAR_DOSSIER.md; G3_SLOT_MATRIX.md.
Our course: grade3/00_course_architecture.md (strands, 42 topics, atom counts, strand tests after Topics 15, 21, 31, 42); the atom registers in /Users/jamesmoore/Documents/Science Knowledge Graph/courses/g35-science/registers/ (the standards codes); CLAUDE.md; reviews/james/G3_ES_LIFE_FEEDBACK_2026-09-30.md §"In chat".
The platform: ~/Documents/CourseBuilding-Vault/40 Platform/TimeBack Scroll takes one row per element, atoms of parts, and a protected test bank.md (six question types; no numeric or text-entry type; PP100 bank served at random; first answer final); build_tools_export_timeback.py (typed items export as AI-graded frq with an accept list; no test element type).
The vault: 50 Decisions/PP100 mastery tests and lesson checks are different jobs.md; 10 Rules/End of Topic Tests assess their own topic.md; 30 QA Methods/An examiner pass checks every test against the released exam.md; 10 Rules/A test item is a fresh instance, never a lesson check reused.md.
Not on disk (said plainly): the released STAAR Grade 5, Florida SSA Grade 5 and AZSCI Grade 5 forms (standards PDFs only); whether Scroll can serve a fixed form, lock Part B to Part A, render multi-select, hot-text or a graph build, offer read-aloud, or expose answer data; any measure of Scroll's AI marker against human marking; any Grade 3 reading-speed figure of our own.
Arithmetic
- Strata allocations: items = round(30 × atoms ÷ strand atoms), floor 1, adjusted to sum 30. Physics 64 atoms: 4, 3, 2, 1, 2, 1, 2, 1, 2, 1, 2, 4, 2, 2, 1 = 30. Matter 29 (22 + 7 SR) over 6 topics of 5, 5, 4, 4, 6, 7 atoms: 5, 5, 4, 4, 6, 6 = 30. Earth & Space 45 over 10 topics of 6, 5, 6, 3, 4, 5, 4, 3, 4, 5: 4, 4, 4, 2, 3, 3, 3, 2, 3, 2 = 30. Life 43 over 11 topics of 5, 4, 3, 6, 3, 4, 3, 3, 5, 3, 4: 3, 3, 2, 4, 2, 3, 2, 2, 3, 2, 4 = 30.
- DOK on 30: up to 10 / about 70 / about 20 per cent → 0 to 3 / 20 to 22 / 6 to 8.
- Bar: 27 of 30 = 90.0 per cent.
- Forms: three attempts, each unseen → 3 fixed forms = 90 protected items per gate (360 for four gates; 450 if Physics is two gates). Practice pool about 60 per gate.
- Time: 28 selected or typed items at about a minute plus two sentences at about two minutes ≈ 32 minutes for a ready child.
- Chance under the retake policy (Becky's Principle 9 arithmetic): if an unready child has a 10 per cent chance of passing one form, three attempts give 27 per cent; instruction between attempts is what keeps that honest, and the per-stratum floor lowers it further.