Do Teacher Expectations Raise Students’ IQ?
A living meta-analysis of the Pygmalion effect — 18 experiments represented by 19 effect-size records, recomputed in your browser.
Cross-examine the headline in 30 seconds
A paper gives an agent a claim. This page also gives it the exact rule, the data, and a way to rerun the calculation.
Run the check yourself, or ask your WebMCP agent to do it. Both update this page.
Buttons are human actions, not an agent simulation.A passed or failed rule is not a scientific truth rating. The historical fit uses 19 records from 18 experiments and does not model their shared-experiment covariance.
Source checks and limitations — current evidence base
Try the real WebMCP agent workflow
Open this page in ChatGPT’s in-app browser, or Chrome with WebMCP enabled. Give your agent this prompt:
Native Chrome test and reproduction instructions · Manual Tool console fallback
Author a document · Import a sample evidence package · Explore the evidence map
If you are a human
Just read. Highlighted sentences are claims — testable assertions with a deterministic check behind them. Ask your agent to challenge one (“is the publication-bias claim actually solid?”), or drive the same tools yourself from the Tool console. When an agent proposes changing the evidence base, nothing happens until you approve it.
If you are an agent
Call get_document_overview first. Every number you cite from this document must come from its
tools, not from parsing the prose. The page computes; its registered rule classifies; you and the human interpret.
Run evaluate_claim; a rule-outcome badge appears in the text. You may propose_study, but only the human can approve it.
Background: the most famous classroom experiment ever run
In 1968, Rosenthal and Jacobson told elementary-school teachers that a test had identified certain pupils as “intellectual bloomers” about to surge ahead. In reality the bloomers had been picked at random. At the end of the year, the researchers reported, those randomly blessed children had gained more IQ than their classmates. The study — Pygmalion in the Classroom — became one of the most cited results in psychology, and its one-line moral, that simply raising a teacher’s expectations makes children measurably smarter, entered textbooks, teacher training, and popular culture.
What entered the textbooks less often: replications began almost immediately, and most of them failed. By 1984, Raudenbush had assembled 18 experiments using the same broad design — deceive teachers about randomly chosen pupils, then measure IQ. This page carries 19 effect-size records because the Pellegrini & Hicks experiment contributes separate aware-tester and blind-tester records.
The evidence base
Across 19 effect-size records from 18 experiments, the pooled random-effects estimate of the expectancy effect is … standard deviations, 95% CI …, p = …. Between-study heterogeneity is moderate (I² = …, Q = …, p = …).
Effect-size records (the table your agent reads through get_studies)
| Record | Experiment | Year | Prior contact | Tester | SMD | Var | Source status | RoB |
|---|
Traceability status: all 19 yi/vi records are secondary-dataset transcriptions with row locators; primary reports checked 0/19; effect-size derivations checked 0/19; structured risk of bias assessed 0/19. The page reproduces calculations over these values; it does not validate their extraction or study design. Source: open metadat distribution; synthesis DOI 10.1037/0022-0663.76.1.85.
Unit warning: the reference fit reproduces the historical 19-row analysis and does not model covariance between the two Pellegrini & Hicks condition records.
What the evidence actually says
Read as one undifferentiated pile, the literature is deflating: pooled across all effect-size records, the average expectancy effect is small and not statistically significant. If the Pygmalion effect were the robust, general phenomenon of popular telling, this is not the forest plot it would leave behind.
But the pile is not undifferentiated, and this is where the story turns. The experiments differ in one crucial, almost embarrassing way: how long the teachers had known their pupils before the researchers lied to them. You cannot easily plant a false expectation about a child the teacher has taught for months. The length of prior teacher–pupil contact is associated, under the fitted capped-linear model, with essentially all of the between-study differences — and once you look only at the studies where the deception had room to work, the effect reappears: in studies where teachers had known their pupils for at most one week, the expectancy induction produced a significant IQ gain.
A narrow descriptive summary is neither “Pygmalion is real” nor “Pygmalion failed to replicate.” In this historical, row-wise synthesis, larger estimates were associated with at most one week of prior contact under the authored subgroup and capped-linear analyses. This does not establish a causal window.
How solid is this?
Two robustness checks are built into this document. First, no single effect-size record changes whether the pooled estimate crosses p < 0.05 — this is a narrow leave-one-record-out threshold check, not a leave-one-experiment-out analysis or proof that the result is stable in magnitude. Second, the evidence base shows no signs of publication bias — though your agent may have opinions about how confidently that can be said. Every claim above carries a document-registered rule; none has run yet on your copy. A badge reports only whether that authored rule passed, failed, or was inconclusive — never scientific truth, validity, risk of bias, or evidence quality. That is deliberate. Don’t take this article’s word for its own claims — have them tested.
PDF vs WebMCP: a testable comparison
No runs recorded — this page makes no claim that WebMCP outperforms PDF. Use the same model, settings and prompt in fresh sessions, then paste the two JSON answers below. Scoring happens only in this browser and never enters the scientific audit ledger.
Download the frozen PDF baseline · read the protocol
Common prompt for both conditions
Ordinary PDF + agent
Attach only the frozen PDF in a fresh session.
Living Evidence + WebMCP agent
Open this page in a fresh WebMCP-enabled session.
Reader’s Workbench
Nothing here yet. Analyses run by your agent (or by you, from the Tool console) will render their figures here, newest first — forest plots, sensitivity analyses, funnel plots, meta-regressions.
Proposed changes to the evidence base
Agents can propose adding studies (for example, experiments published after this document was written). Nothing enters the evidence base without a human pressing Approve.
Audit ledger
A device-persistent SHA-256 hash chain of analyses, registered-rule outcomes and decisions. It survives reloads in this browser. The chain detects edits and reordering; it is not a trusted timestamp or author identity.
Reproducibility receipt
Create a signed receipt for the current scientific-state hash and audit-chain head. The signing key is generated for this page load and rotates on reload; each receipt is individually self-signed. A published receipt fixes only the audit prefix it covers, while later reader actions form a local suffix. Save the detached artifact receipt and pin its fingerprint in an external archive or repository release before using it as authorship evidence.
Tool console
Drive the document’s tools by hand (no agent required)
This is the same tool surface an AI agent sees through WebMCP — same schemas, same effects, same ledger. It exists so that nothing an agent can do to this document is hidden from you.
Methods note
Pooling uses a random-effects model with REML estimation of τ² (DerSimonian–Laird and fixed-effect
available through the tools); heterogeneity is reported as Cochran’s Q and I²; the moderator model is a mixed-effects
meta-regression on prior contact truncated at three weeks, following Raudenbush (1984); publication-bias diagnostics
use Egger’s regression test. The engine is dependency-free JavaScript running in this page; selected numerical
outputs for this fixture reproduce R metafor reference values to the tested precision. That is a software
reproduction check, not validation of the extracted data, model assumptions, or scientific conclusion. Statistical caveat: subgroup and moderator analyses here are observational
comparisons across randomized experiments and can be confounded by other study features.
References
Rosenthal, R., & Jacobson, L. (1968). Pygmalion in the classroom: Teacher expectation and pupils’ intellectual development. Holt, Rinehart & Winston.
Raudenbush, S. W. (1984). Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction: A synthesis of findings from 18 experiments. Journal of Educational Psychology, 76(1), 85–97.
Raudenbush, S. W., & Bryk, A. S. (1985). Empirical Bayes meta-analysis. Journal of Educational Statistics, 10(2), 75–98.
Viechtbauer, W. (2010). Conducting meta-analyses in R with the metafor package. Journal of Statistical Software, 36(3), 1–48.