Do Teacher Expectations Raise Students’ IQ?

A living meta-analysis of the Pygmalion effect — 18 experiments represented by 19 effect-size records, recomputed in your browser.

Initializing agent interface…

Cross-examine the headline in 30 seconds

A paper gives an agent a claim. This page also gives it the exact rule, the data, and a way to rerun the calculation.

See the shared workbench
Not run on this device

Run the check yourself, or ask your WebMCP agent to do it. Both update this page.

Buttons are human actions, not an agent simulation.

A passed or failed rule is not a scientific truth rating. The historical fit uses 19 records from 18 experiments and does not model their shared-experiment covariance.

Try the real WebMCP agent workflow

Open this page in ChatGPT’s in-app browser, or Chrome with WebMCP enabled. Give your agent this prompt:

Native Chrome test and reproduction instructions · Manual Tool console fallback

Author a document · Import a sample evidence package · Explore the evidence map

This is a Living Evidence document. The statistics in this article are not typeset into the text — they are computed live, in your browser, from the effect-size records embedded in this page. If you are reading with an AI agent in a WebMCP-enabled browser, your agent has been handed tools to re-run, stress-test, and extend every analysis here. Whatever it does is rendered back into this document — in the Reader’s Workbench and the audit ledger — so you watch the cross-examination happen on the page itself.

If you are a human

Just read. Highlighted sentences are claims — testable assertions with a deterministic check behind them. Ask your agent to challenge one (“is the publication-bias claim actually solid?”), or drive the same tools yourself from the Tool console. When an agent proposes changing the evidence base, nothing happens until you approve it.

If you are an agent

Call get_document_overview first. Every number you cite from this document must come from its tools, not from parsing the prose. The page computes; its registered rule classifies; you and the human interpret. Run evaluate_claim; a rule-outcome badge appears in the text. You may propose_study, but only the human can approve it.

Background: the most famous classroom experiment ever run

In 1968, Rosenthal and Jacobson told elementary-school teachers that a test had identified certain pupils as “intellectual bloomers” about to surge ahead. In reality the bloomers had been picked at random. At the end of the year, the researchers reported, those randomly blessed children had gained more IQ than their classmates. The study — Pygmalion in the Classroom — became one of the most cited results in psychology, and its one-line moral, that simply raising a teacher’s expectations makes children measurably smarter, entered textbooks, teacher training, and popular culture.

What entered the textbooks less often: replications began almost immediately, and most of them failed. By 1984, Raudenbush had assembled 18 experiments using the same broad design — deceive teachers about randomly chosen pupils, then measure IQ. This page carries 19 effect-size records because the Pellegrini & Hicks experiment contributes separate aware-tester and blind-tester records.

The evidence base

Across 19 effect-size records from 18 experiments, the pooled random-effects estimate of the expectancy effect is standard deviations, 95% CI , p = . Between-study heterogeneity is moderate (I² = , Q = , p = ).

Forest plot of the current evidence base. Squares are study estimates sized by weight; the diamond is the pooled random-effects (REML) estimate. This figure re-renders whenever the evidence base changes.
Effect-size records (the table your agent reads through get_studies)
RecordExperimentYearPrior contactTesterSMDVarSource statusRoB

Traceability status: all 19 yi/vi records are secondary-dataset transcriptions with row locators; primary reports checked 0/19; effect-size derivations checked 0/19; structured risk of bias assessed 0/19. The page reproduces calculations over these values; it does not validate their extraction or study design. Source: open metadat distribution; synthesis DOI 10.1037/0022-0663.76.1.85.

Unit warning: the reference fit reproduces the historical 19-row analysis and does not model covariance between the two Pellegrini & Hicks condition records.

What the evidence actually says

Read as one undifferentiated pile, the literature is deflating: pooled across all effect-size records, the average expectancy effect is small and not statistically significant. If the Pygmalion effect were the robust, general phenomenon of popular telling, this is not the forest plot it would leave behind.

But the pile is not undifferentiated, and this is where the story turns. The experiments differ in one crucial, almost embarrassing way: how long the teachers had known their pupils before the researchers lied to them. You cannot easily plant a false expectation about a child the teacher has taught for months. The length of prior teacher–pupil contact is associated, under the fitted capped-linear model, with essentially all of the between-study differences — and once you look only at the studies where the deception had room to work, the effect reappears: in studies where teachers had known their pupils for at most one week, the expectancy induction produced a significant IQ gain.

A narrow descriptive summary is neither “Pygmalion is real” nor “Pygmalion failed to replicate.” In this historical, row-wise synthesis, larger estimates were associated with at most one week of prior contact under the authored subgroup and capped-linear analyses. This does not establish a causal window.

How solid is this?

Two robustness checks are built into this document. First, no single effect-size record changes whether the pooled estimate crosses p < 0.05 — this is a narrow leave-one-record-out threshold check, not a leave-one-experiment-out analysis or proof that the result is stable in magnitude. Second, the evidence base shows no signs of publication bias — though your agent may have opinions about how confidently that can be said. Every claim above carries a document-registered rule; none has run yet on your copy. A badge reports only whether that authored rule passed, failed, or was inconclusive — never scientific truth, validity, risk of bias, or evidence quality. That is deliberate. Don’t take this article’s word for its own claims — have them tested.

PDF vs WebMCP: a testable comparison

No runs recorded — this page makes no claim that WebMCP outperforms PDF. Use the same model, settings and prompt in fresh sessions, then paste the two JSON answers below. Scoring happens only in this browser and never enters the scientific audit ledger.

Download the frozen PDF baseline · read the protocol

Baseline 2026-09-04.v1 · SHA-256 loading…. The PDF and this page share the same 19 numeric records. This evaluates three specified tasks on one corpus; it does not establish general accuracy, speed, or scientific validity.

Common prompt for both conditions

      
    

Ordinary PDF + agent

Attach only the frozen PDF in a fresh session.

Living Evidence + WebMCP agent

Open this page in a fresh WebMCP-enabled session.

Reader’s Workbench

Nothing here yet. Analyses run by your agent (or by you, from the Tool console) will render their figures here, newest first — forest plots, sensitivity analyses, funnel plots, meta-regressions.

Proposed changes to the evidence base

Agents can propose adding studies (for example, experiments published after this document was written). Nothing enters the evidence base without a human pressing Approve.

Audit ledger

A device-persistent SHA-256 hash chain of analyses, registered-rule outcomes and decisions. It survives reloads in this browser. The chain detects edits and reordering; it is not a trusted timestamp or author identity.

    Reproducibility receipt

    Create a signed receipt for the current scientific-state hash and audit-chain head. The signing key is generated for this page load and rotates on reload; each receipt is individually self-signed. A published receipt fixes only the audit prefix it covers, while later reader actions form a local suffix. Save the detached artifact receipt and pin its fingerprint in an external archive or repository release before using it as authorship evidence.

    Tool console

    Drive the document’s tools by hand (no agent required)

    This is the same tool surface an AI agent sees through WebMCP — same schemas, same effects, same ledger. It exists so that nothing an agent can do to this document is hidden from you.

    Methods note

    Pooling uses a random-effects model with REML estimation of τ² (DerSimonian–Laird and fixed-effect available through the tools); heterogeneity is reported as Cochran’s Q and I²; the moderator model is a mixed-effects meta-regression on prior contact truncated at three weeks, following Raudenbush (1984); publication-bias diagnostics use Egger’s regression test. The engine is dependency-free JavaScript running in this page; selected numerical outputs for this fixture reproduce R metafor reference values to the tested precision. That is a software reproduction check, not validation of the extracted data, model assumptions, or scientific conclusion. Statistical caveat: subgroup and moderator analyses here are observational comparisons across randomized experiments and can be confounded by other study features.

    References

    Rosenthal, R., & Jacobson, L. (1968). Pygmalion in the classroom: Teacher expectation and pupils’ intellectual development. Holt, Rinehart & Winston.

    Raudenbush, S. W. (1984). Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction: A synthesis of findings from 18 experiments. Journal of Educational Psychology, 76(1), 85–97.

    Raudenbush, S. W., & Bryk, A. S. (1985). Empirical Bayes meta-analysis. Journal of Educational Statistics, 10(2), 75–98.

    Viechtbauer, W. (2010). Conducting meta-analyses in R with the metafor package. Journal of Statistical Software, 36(3), 1–48.