Every TOEFL Speaking task — independent or integrated — is graded against the same rubric, on a 0–4 scale per task. Raters (or AI-assisted scoring calibrated against human raters) aren't listening for a native accent or flawless grammar. They're scoring three specific things: Delivery, Language Use, and Topic Development. Knowing what each one actually means changes how you practice.
The Three Criteria
Delivery
How clear and fluid your speech is — pace, pronunciation, and intonation. This is not about sounding American or British; it's about whether a listener has to work to understand you. Long pauses, choppy phrasing, or speech that's too fast to follow all cost points here, even if every word is grammatically correct.
Language Use
The range and accuracy of the grammar and vocabulary you control. A response built entirely from short, simple sentences caps your score here even if every sentence is correct — raters are also listening for whether you can handle more complex structures (relative clauses, conditionals, varied connectors) without falling apart.
Topic Development
Whether your response is complete, coherent, and directly answers the prompt. For the integrated tasks (3 and 4) this specifically means: did you accurately report what the reading and the audio said, not just your personal reaction to it. A beautifully delivered response that skips half the source content scores low here.
The 0–4 Scale, By Band
| Score | What it sounds like |
|---|---|
| 4 | Sustained, coherent, well-paced. Minor lapses in grammar or word choice don't get in the way of understanding. |
| 3 | Generally clear and coherent, but with noticeable hesitation, some vague or imprecise language, or a minor gap in content. |
| 2 | Response is understandable but effortful — limited development, noticeable pauses, or grammar/vocabulary errors that occasionally obscure meaning. |
| 1 | Very limited content or largely unintelligible — the response barely addresses the task. |
| 0 | No attempt, or a response entirely unrelated to the task. |
The gap that matters most
Most test-takers plateau at a 2 because of Topic Development, not Delivery — they run out of things to say before time's up, or (on Tasks 3–4) they give their own opinion instead of reporting the reading and lecture. Practicing what to say is usually higher-leverage than practicing how to say it.
How This Differs Across the Four Tasks
- ·Tasks 1–2 (Independent, 45s response): Topic Development means picking a clear position and supporting it with one well-explained reason — not three shallow ones. Raters reward depth over breadth here.
- ·Task 3 (Integrated Campus, 60s response): you must report BOTH the announcement and the student's opinion about it. A response that only summarizes one side is incomplete regardless of how well it's delivered.
- ·Task 4 (Integrated Academic, 60s response): you must connect the reading's definition to the professor's example(s). Dropping the connection — stating the definition and the example as two unrelated facts — is the single most common Topic Development error on this task.
A Fast Way to Self-Check
Record a response, then play it back and ask three questions in order: Did I address every part of the prompt (Topic Development)? Could a stranger follow it without replaying (Delivery)? Did I use anything beyond simple sentences (Language Use)? Whichever question you hesitate on is the one to drill next, not the one that feels most uncomfortable.
Practice against the real rubric
Our Speaking section times every task exactly like the real exam and scores your response, plus AI feedback that flags pacing and filler words against this same rubric.
Go to Speaking practiceFor sentence-by-sentence structures to fill in under this rubric, pair this with our TOEFL Speaking Templates guide — the rubric tells you what's being measured, the templates give you a repeatable way to hit it under time pressure.