← All topics (review home)

Skill: write-like-james → write-mcqs, review-questions (rules read; nothing authored) @ 84b31cb

Grade 3 Science end-of-strand tests: a format proposal for external validation

From James Moore's Grade 3 Science build, 30 September 2026 (revised the same evening against Becky's "A - Mastery Gates" and "X - Blueprint Creation and Form Review"). For Becky (Alpha assessment) and James. Nothing in this document is authored test content; it is the format we propose to build, set beside David's Alpha Standardized Science Grade 3 blueprint and Becky's design rules, with the questions we need answered before any item is written.

What the course is. Grade 3 Science is one self-paced course on TimeBack Scroll: 183 one-idea lessons in 42 small topics, in four strands taught in order (Physics, Matter, Earth & Space, Life). Every topic ends with a PP100 drawn from a bank of about fifty plain items, one atom (knowledge component) per item. The course follows TEKS elementary science, NGSS 3-5, Florida NGSSS and Arizona, so it teaches content the NGSS Grade 3 band leaves to other grades: matter and its states, energy, light, sound, the Sun, stars and solar system, soil, natural resources, food chains.

What the tests are for. Four end-of-strand tests, one after each strand, decide whether the child is ready to move on to the next strand. In Becky's terms they are within-course instructional gates, not endpoint mastery assessments. The Grade 3 endpoint already exists as David's family of ten live forms (G3.22 to G3.31); we do not build one.

What we ask Becky to validate, and by when. Four things, in this order: (1) the assessment claim and the classification of the strand test as a within-course gate (section "The gate in Becky's terms"); (2) the synthetic blueprint written in her ten-step shape (section "The blueprint"), in particular the pitch, the strata and the retake policy; (3) the eight questions below, which her documents raise or leave open; (4) once she has signed off the format, five form plans for the Physics gate (slots only, no items) and then one authored form, each reviewed against Steps 8 and 9 of her process before the other forms are written. We propose the format and claim this week and the Physics form plans within a week of her sign-off; the dates are hers and James's to set. Nothing goes live on her side of the gate.


Questions

1. Is 30 items the right length for a fixed-form strand gate, with the cut of 27 chosen alongside it?

David's number; three-quarters of the 36-item Grade 3 endpoint form. Becky's design rule 4 says length and cut are chosen together for a fixed form; her Economics example is 30 items, 27 to pass, about 30 minutes, untimed.

Recommendation: 30 items, untimed, 27 to pass, a ready child finishing in about 30 minutes. Not an hour. Choices: 30 / 24 (shorter, same cut ratio) / other.

2. Should Physics, with 64 atoms in 15 topics, be one gate or two?

Becky's rule 2 sizes a within-course gate at about a dozen knowledge components; her Issue 1 says a hundred is an exam, not a gate. The topic PP100s already gate 3 to 8 atoms each; the strand test samples across 22 (Matter) to 64 (Physics) atoms.

Recommendation: one gate per strand, sampling by topic strata, with Physics split into two gates (Motion and Forces, Topics 1 to 10; Sound, Energy and Light, Topics 11 to 15) if Becky judges 64 too broad to gate at 30 items. Choices: one Physics gate / two.

3. Besides 27 of 30, does the gate require at least one correct item in every topic stratum?

Becky's rule 4: if mastery requires several essential facets, set a requirement for each rather than averaging them away. A child could score 27 by missing all three items of one topic.

Recommendation: yes: 27 of 30 and no topic stratum at zero correct; a stratum at zero routes the child to that topic's lessons and PP100, whatever the total. Choices: yes / total only / a per-stratum floor only for prerequisite-heavy topics.

4. Is the retake policy right: at most two retests, each on an unseen fixed form after the missed topics' instruction, and what happens after a third fail?

Becky's Principle 9: unlimited retakes destroy the cut; failure must change the route; the pass rule is validated across the whole retake policy. Principle 6: a retest shows only unseen items.

Recommendation: three attempts in all (Forms A, B, C); after a third fail the child redoes the failed topics' lessons and PP100s in full and a teacher decides. Choices: as recommended / two attempts / other.

5. Which fail scores route the child straight back to instruction, and which send them through the topic PP100s first?

Becky's rule 5 asks for both bands to be named before launch.

Recommendation: below 21 of 30 (70 per cent): back to the lessons of every topic with a missed item; 21 to 26: redo the PP100 of each topic with a missed item, then the next form. Choices: as recommended / other bands.

6. May one-sentence constructed items sit in the gate before Scroll's marking has been checked against human marking?

Becky's Principle 8 says free response is often needed because explaining is part of the claim, with a binary mark and an excessively clear stem; marker reliability is a real error source. Scroll marks free response with an AI grader we have not yet tested on a child's sentence.

Recommendation: two one-sentence items per form, binary-marked, but only after a sample of 50 child answers has been marked by the grader and by James or Olivia and the agreement recorded; until then those two slots are MCQs. Choices: as recommended / never in the gate / in from the start.

7. Which Alpha item types can TimeBack Scroll render, and are they familiar enough from practice not to count as unfamiliar controls?

Scroll's contract on our disk lists MCQ, AI-graded free response, match pairs, cloze, sequence and sort; no multi-select, no drag-to-build graph, no hot-text, no numeric-entry type. Becky's Principle 5 names unfamiliar controls as construct-irrelevant difficulty.

Recommendation: use only Scroll types the child has met in the practice form; map the rest as in Card 3; ask the Scroll team today. Choices: as mapped / wait for Scroll's answer before fixing the mix.

8. Who certifies the Grade 3 content the AlphaTest Grade 3 family leaves out?

David's spec tests only what the old Alpha Grade 3 course taught and excludes matter, energy, light, astronomy, soil and food chains. James's course teaches all of them at Grade 3 (about 80 of 183 atoms).

Recommendation: the strand gates certify readiness on that content now; David re-derives the Grade 3 family's testable-content list from James's course when it is live, under his own Q8/Q9 rule. Choices: as recommended / strand gates only / revise the family first.

Settled by Becky's documents, so not asked (answer and source):

Applied without asking (rulings already given): static Form A (James, today); the 90 per cent bar; Matter gets a gate like the other three; tests carry no hints or frames; every key passes blind review; an end-of-strand test gets the examiner pass against released forms before its blind read; David's untimed rule (Q7) and his independent scoring of two-part items (Q6).


Becky's WorkFlowy guidance: reconciliation

Read in full: reference/becky_WF_A_Mastery_Gates_2026-09-30.md and reference/becky_WF_X_Blueprint_Creation_2026-09-30.md, including the Economics, New York, MAP and Grade 5 Mathematics examples.

Where the proposal already followed her logic: protected, blind-reviewed items; one atom per item; fresh instances; a fixed form with the cut chosen with its length; a strand-scoped claim; a synthetic blueprint informed by external forms; released items as anchors, never copied; live-data review after launch.

Where it conflicted, and what we change:


The proposed format beside David's Grade 3 spec

Cells marked per Becky's A or per Becky's X changed on her documents.

David's Grade 3 endpoint form (G3.22 to G3.31)Proposed end-of-strand gate (Grade 3)
Kind of instrumentcourse/grade endpoint mastery assessmentwithin-course instructional gate on one strand (per Becky's A, Purpose and Principle 10)
Items per form36 nominal (34 to 38)30
Marksabout 41 to 44 (multi-select and two-part 2 pt)30, one binary mark per item; a two-part pair is two items (per Becky's A, Issue 12)
Timeuntimed; about 35 to 50 minutes expecteduntimed; about 30 minutes for a ready child
Scopeall 18 NGSS Grade 3 PEs at least once, plus taught extensionsone strand, sampled by topic strata in proportion to atoms; no atom compulsory or fixed to a slot (per Becky's X, Step 3)
Cumulative samplingnone at Grade 3 (25 to 30 per cent at Grade 5)none
Pitch and DOK mix0 to 2 DOK 1 / 19 to 20 DOK 2 / 14 to 15 DOK 3 (rigour superset)0 to 3 DOK 1 / 20 to 22 DOK 2 / 6 to 8 DOK 3; DOK 3 only where the standard's practice is DOK 3; at least one DOK 2 in every stratum (per Becky's A, Principles 2 and 4, Issue 3)
Clusters7 phenomenon stimuli × 2 to 3 items, about 50 per cent5 to 6 stimuli × 2 to 3 items, about 40 to 50 per cent; a stimulus is one short case with one figure
Reading loadstimulus ≤ 100 to 150 words; sentences ≤ 15 wordsstimulus ≤ 80 words, one figure; sentences ≤ 12 words; one judgement per item (per Becky's A, Principle 5, young children)
Item types15 MC · 5 multi-select · 3 two-part pairs (6) · 3 gap-match · 2 inline choice · 2 hot-text · 1 match or order · 2 numeric entry14 MC · 3 two-part pairs (6) · 3 match or sort · 2 cloze · 1 sequence · 2 numeric entry · 2 one-sentence constructed (Card 3; per Becky's A, Principle 8 and rule 3 for the constructed items)
Free responsenone (Q12)2 one-sentence items, binary-marked, excessively clear stems; MCQs in those slots until the marker check passes (per Becky's A, Principle 8)
Vocabulary6 to 8 vocabulary-bearing DOK 2 items4 to 6, the strand's coined terms inside application items
Complexity tagseasy / medium / hardcomplexity, not difficulty; no numerical difficulty targets before live data (per Becky's A, Issue 3)
Mastery bar89.5 per cent of points, one attempt per form27 of 30, and (proposed) no topic stratum at zero correct (per Becky's A, rule 4; question 3)
Retakesretake on another form; ten forms for volumeat most two retests, each on an unseen form after the missed topics' instruction; fail bands named (per Becky's A, Principles 6 and 9, rule 5)
Forms and bank10 live forms per grade, no stimulus or template reused3 protected fixed forms per gate (A, B, C), planned together; a separate practice pool of about 60 (per Becky's A, Principle 6; X, Step 7)
Sourcingreleased items first, generation lastthe same order, filtered to the strand's atoms; our own items where none exist (most of Matter, energy, light, sound)
Read-aloudyes at Grade 3 (stems, prompts, options)yes if Scroll offers it; note Becky's caution that audio swaps a reading load for a listening-and-memory load
Validationplan audit → authoring → blind solve → cross review → validator → human walkthe same stages in our tools; one form then five varied form plans compared; Becky's gate before live; monitoring indicators named before launch (per Becky's X, Principle 6, Steps 8 to 10)

The gate in Becky's terms

The call: the strand test is a within-course instructional gate on a bundle of 22 to 64 knowledge components; that fixes its pitch, its threshold logic, its stopping rule and its failure routing.

What changes: the gate is written up as above in the strand test's record; the PP100 (about a dozen atoms per topic cluster) stays the fine-grained gate that certifies each atom.


The blueprint (in the shape of Becky's "X - Blueprint Creation", ten steps)

A synthetic blueprint: no single external assessment is the target. David's Grade 3 spec and the released Grade 3 forms (Tennessee TCAP, Louisiana LEAP, Arkansas ATLAS practice) inform the pitch and the item grammar; James's course defines the content.

Step 1: the assessment claim

For Grade 3 Science students who have completed a strand, the gate result is intended to support the conclusion that they can apply that strand's taught ideas to fresh cases at the level the lessons taught, for the purpose of deciding whether they move on to the next strand or return to named topics, in relation to James Moore's Grade 3 Science course (TEKS, NGSS 3-5, Florida, Arizona). It is not intended to support conclusions about mastery of every atom (the topic PP100s do that), about readiness for the Grade 3 endpoint AlphaTest, about performance on any state test, or about hands-on investigation skill.

Step 2: external assessments and how they inform the blueprint

Similarity recorded qualitatively: closest to LEAP in shape, to TCAP in length, to none in content (the strand scope is ours).

Step 3: content coverage model

Step 4: item mix and cognitive demand

Step 5: acceptable item boundaries (task sketches, not items)

Bands as in Becky's examples: too low / on target lower boundary (DOK 1) / on target secure (DOK 2) / on target upper boundary (DOK 3) / too high. One stratum per strand here; the rest follow after her sign-off of the format, before any form plan.

Step 6: whole-assessment requirements

Step 7: production and assembly

Step 8 and Step 9: one form, then several

Step 10: monitoring from the first live sitting


Card 3: item grammar, and what Scroll must confirm

The call: build the Alpha item grammar on the six types Scroll renders, introduce every type in the practice form first, and ask the Scroll team about the rest today.

Alpha type (per 36)On Scroll (per 30)How it rendersConfirm with Scroll
MC single, 4 options (15)14 MCmcq, four optionsnone
Multi-select, 5 options, 2 pt (5)0; the demand moves into two-part pairs and sort itemsno multi-select typea select-all type
Two-part evidence pairs (3 pairs)3 pairs = 6 itemstwo mcq rows sharing one stimulus, repeated on each row; scored independentlywhether Part B can be locked to Part A
Gap-match incl. graph construction (3)3 match or sort (match_pairs, sort_categories)no drag-to-build grapha gap-match or graph type
Inline choice (2)2 cloze with a word bankclozenone
Hot-text (2)0; asked as an MC whose options are the sentencesno hot-text typea hot-text type
Match or order (1)1 sequencesequence, never presortednone
Numeric entry (2)2 numeric entry as AI-graded frq with an exact accept listno numeric typea numeric type; whether the grader honours an exact list
Constructed response (0 at Grade 3)2 one-sentence items, binary (question 6)AI-graded frq with one-point marking guidancemarker agreement check first

Why:

What changes: the export tool needs a fixed-form gate and a protected gate bank beside the PP100 bank; both are contract questions for the Scroll team, listed in the placeholders file.


What changes for the course

What sits at the end of each strand to prepare the child

The call: a practice form from its own pool and two free revision pages; nothing else now.

ResourceWhat it isCostRecommendation
Practice formone drawn form in the gate's grammar, from a separate practice pool of about 60, a feedback line under every itemone pool's authoring per gatebuild, after the gate; every gate item type is met here first (per Becky's A, Principles 5 and 6)
Topic rule sheetsone page per topic: each lesson's rule sentence in teaching order, built from the lesson filesnone; the renderer already builds this blockbuild now
Summary-video playlistthe strand's topic summary videos in ordernonebuild now
Key-terms typing settyped recall over the strand's coined termsa day's authoring per strandhold: the PP100 mixed practice already recalls terms
Missed-item replaythe child's wrong items served again with feedbackneeds answer data from the platformask the Scroll team; note it may only use practice-pool items, never gate items (per Becky's A, Principle 6)

Placeholders

grade3/strand_tests/PLACEHOLDERS.md lists the four gates (Physics, Matter, Earth & Space, Life) with status "format under validation with Becky", their position in the course tree (after Topics 15, 21, 31 and 42), and the intended values of the ledger rows the TimeBack hand-off should carry.

What happens next

James forwards this page to Becky.

Becky and James answer the eight questions; her verdict on the claim, the pitch and the retake policy comes first.

The Scroll team answers Card 3's rendering questions.

Then five Physics form plans (slots only), reviewed against Steps 8 and 9, then one authored form.


David's message to James, and Becky's words, verbatim (30 September 2026)

David: "30 questions per end-of-domain test feels right — about three-quarters of a full end-of-grade form (our G3 forms run 36 items; middle school 40). The thing I'd hold onto isn't the length, it's the item grammar. Build them from the same blueprint DNA as the Alpha tests: phenomenon-driven clusters (the passage-type experience), at least one numeric-entry item, drag-and-drop/matching, two-part evidence items, and a DOK mix weighted toward 2s and 3s; keep the 90% mastery bar so 'passed' means the same thing at every rung. The end-of-unit layer becomes the format on-ramp — leave clear simple MCQs for initial mastery and move the format-preparation job to these tests, and let the AlphaTest be pure validation. The end-of-year Grade 3 test already exists as a family: ten live G3 forms (G3.22–31) built to one spec from released items (STAAR, TN TCAP, AR ATLAS, NY ILS, KY KSA, WI Forward, LA LEAP, GA Milestones, NAEP, TIMSS; NWEA MAP percentile anchor). You can change this spec for the end-of-domain tests or use it as-is."

Becky, relayed: "It is very important you don't just write something and expect to change it in the future. External validation is critical. And once a mastery test is in place it is extremely difficult to change."

James's ask, in chat (30 September, about 15:15): end-of-strand tests for Physics, Earth & Space, Life (static; 90 per cent on ours means 90 per cent on any equivalent state test; a large bank for retests; some text entry and short free response beside MCQ); a plan for length and item count ("an hour?"); a proposal for an end-of-strand practice test and revision resources.

Evidence on disk, and what is not

Becky's pages (exports): reference/becky_WF_A_Mastery_Gates_2026-09-30.md (Principles 1 to 10, design rules 1 to 7, the PP100 revision Issues 1 to 12); reference/becky_WF_X_Blueprint_Creation_2026-09-30.md (Principles 1 to 9, Steps 1 to 10, the Economics, New York, MAP and Grade 5 Mathematics examples).

David's blueprint folder, reference/alpha_test_blueprints/: G3_TEST_DESIGN_SPEC_v1.md (36-item form; 18 PEs; DOK ≤5 / ~55 / ~40; 7 clusters; item-type table; 89.5 per cent bar; Q1 numeric entry only, Q6, Q7 untimed, Q8/Q9 testable content, Q12 no FRQ; 2026-08-18 read-aloud amendment); G3_RESEARCH_DOSSIER_2026-08-07.md (TCAP 25 to 26 items, 50 minutes; LEAP 36 items, two 70-minute sessions; ATLAS 6 clusters + 10, about 75 minutes; MAP 40 to 43 items, median 32 to 37 minutes; Iowa 30 items in 35 minutes; NAEP about 90 seconds an item; Georgia DOK bands); G4_G5_DECISION_CHECKPOINT_2026-08-10.md (Grade 5 cumulative band); G6G8_DECISION_CHECKPOINT_2026-08-11.md (40-item forms; false-fail arithmetic); build_refs_Item_Writing_Rules.md; EXEMPLAR_DOSSIER.md; G3_SLOT_MATRIX.md.

Our course: grade3/00_course_architecture.md (strands, 42 topics, atom counts, strand tests after Topics 15, 21, 31, 42); the atom registers in /Users/jamesmoore/Documents/Science Knowledge Graph/courses/g35-science/registers/ (the standards codes); CLAUDE.md; reviews/james/G3_ES_LIFE_FEEDBACK_2026-09-30.md §"In chat".

The platform: ~/Documents/CourseBuilding-Vault/40 Platform/TimeBack Scroll takes one row per element, atoms of parts, and a protected test bank.md (six question types; no numeric or text-entry type; PP100 bank served at random; first answer final); build_tools_export_timeback.py (typed items export as AI-graded frq with an accept list; no test element type).

The vault: 50 Decisions/PP100 mastery tests and lesson checks are different jobs.md; 10 Rules/End of Topic Tests assess their own topic.md; 30 QA Methods/An examiner pass checks every test against the released exam.md; 10 Rules/A test item is a fresh instance, never a lesson check reused.md.

Not on disk (said plainly): the released STAAR Grade 5, Florida SSA Grade 5 and AZSCI Grade 5 forms (standards PDFs only); whether Scroll can serve a fixed form, lock Part B to Part A, render multi-select, hot-text or a graph build, offer read-aloud, or expose answer data; any measure of Scroll's AI marker against human marking; any Grade 3 reading-speed figure of our own.

Arithmetic