Germany’s immigrant gap is not a tracking artefact
The most common explanation for the gap between immigrant and native-born students in Germany is that immigrant children are sorted into the lower tracks. Holding both socio-economic status and track constant leaves most of the gap standing.
Summary
Germany’s PISA 2022 results were the lowest the country has recorded in all three subjects, and the debate that followed reached immediately for two explanations: socio-economic disadvantage and early tracking. Both are real and both are measurable in the public microdata. Neither, it turns out, accounts for most of the immigrant gap.
This note takes the explanation seriously enough to test it.
What the data shows.
- First-generation immigrant students scored 96.6 points below native-born students (6.2). Holding socio-economic status constant leaves 66.2 points. Holding school track constant as well still leaves 53.9 points (5.4).
- So of the 96.6-point raw gap, socio-economic composition accounts for roughly a third and track placement for roughly a further eighth. More than half survives both.
- Tracking itself is nonetheless the largest single structure in the data. Students in the basic general track scored 149.9 points below Gymnasium students (5.9), and 130.1 points below once socio-economic status is held constant. That is roughly one and a half national standard deviations.
- Germany’s socio-economic gradient is the steepest of the eight systems in this series: 39.6 score points per unit of the PISA index (1.5), against a published OECD figure of 19% of achievement variance accounted for.
The distinction matters because the two explanations imply different policies. If the immigrant gap were mostly track placement, the lever would be the transition decision at the end of primary school. Since it is mostly not, the lever is whatever happens to immigrant students inside the track they are in, and language support is the obvious candidate this note cannot test directly.
Why this is the question for Germany
Germany is the country where the tracking hypothesis is most plausible and most often asserted. Children are sorted at about age ten into tracks with sharply different curricula and destinations, immigrant children are over-represented in the lower tracks, and the sorting happens early enough that a language disadvantage at age ten becomes a structural placement for the following decade. The argument is coherent, which is exactly why it deserves a number rather than agreement.
The number is obtainable because Germany’s public file identifies the track. It does not do so in the obvious place: the ISCED programme code records every German 15-year-old as lower secondary general, and the stratum variable is suppressed entirely for Germany in the public file. The national study-programme code carries it, with published labels naming the track structure directly. An analyst who checks only the two standard variables concludes that Germany cannot be analysed by track, which is wrong.
The decomposition below is therefore the whole point of the note: raw gap, then net of socio-economic status, then net of socio-economic status and track together.
The data, and how it was handled
6,116 students in 257 schools, representing a weighted population of 681,399 15-year-olds, from the OECD’s PISA 2022 student and school Public Use Files. National mean in mathematics: 474.8 points (SE 3.1).
Achievement is imputed rather than measured, so estimates are pooled by Rubin’s rules over all ten plausible values. The sample is clustered in schools, so standard errors come from the 80 Fay-adjusted replicate weights the OECD ships for the purpose. Omitting either understates the uncertainty, usually by a factor of two or more.
Before presenting anything new, this note reproduces figures the OECD has already published for this country.
| Published by the OECD | Published | This analysis | SE |
|---|---|---|---|
| Advantaged minus disadvantaged, mathematics | 111 | 111.61 | 4.74 |
| Share of mathematics variance accounted for by ESCS | 0.19 | 18.7% | 0.013 |
| Girls minus boys, reading | 19 | 19.41 | 3.07 |
Source: OECD, PISA 2022 Results (Volume I and II) Country Note: Germany, published 5 December 2023, https://www.oecd.org/en/publications/pisa-2022-results-volume-i-and-ii-country-notes_ed6fbcc5-en/germany_1a2cf137-en.html (read 6 August 2026)
The immigrant gap, and what survives each adjustment
4,030 native-born students, 903 second-generation and 474 first-generation. The socio-economic difference between them is real but modest; the achievement difference is not modest at all.
| Immigrant background | Students | Share | Mean, mathematics | SE | Mean ESCS |
|---|---|---|---|---|---|
| Native | 4,030 | 65.4% | 495.0 | 3.0 | 0.05 |
| Second generation | 903 | 14.7% | 457.2 | 4.3 | -0.60 |
| First generation | 474 | 8.1% | 398.4 | 5.9 | -0.78 |

Socio-economic status takes a substantial bite out of both gaps, as it should in the OECD’s steepest gradient. It does not take most of it.
| Contrast | Unadjusted | SE | Net of ESCS | SE |
|---|---|---|---|---|
| Second generation minus Native | -37.8 | 4.5 | -14.0 | 4.1 |
| First generation minus Native | -96.6 | 6.2 | -66.2 | 6.2 |
Adding track to the adjustment is the test the tracking hypothesis asked for. For second-generation students the gap net of socio-economic status and track is 17.9 points (3.4); for first-generation students, 53.9 points (5.4). Comparing like with like inside the same track and at the same socio-economic level, the gap is smaller than the headline but is emphatically still there.
One reading deserves to be blocked before it starts. This is not evidence that tracking is harmless. It is evidence that tracking does not explain the immigrant gap. Those are different claims, and the second says nothing about the first.
The track gap, for scale
Since the note has built the track variable, it is worth reporting on its own, because the numbers are larger than anything else in Germany’s data.
| School track | Students | Share | Mean, mathematics | SE | Mean ESCS |
|---|---|---|---|---|---|
| Academic track (Gymnasium) | 2,270 | 36.5% | 546.5 | 2.6 | 0.38 |
| Comprehensive school | 1,379 | 23.2% | 435.3 | 4.8 | -0.39 |
| Intermediate general track | 1,551 | 25.0% | 455.5 | 3.5 | -0.40 |
| Basic general track | 617 | 10.0% | 396.7 | 5.2 | -0.75 |

| Contrast | Unadjusted | SE | Net of ESCS | SE |
|---|---|---|---|---|
| Comprehensive school minus Academic track (Gymnasium) | -111.2 | 5.4 | -94.6 | 5.7 |
| Intermediate general track minus Academic track (Gymnasium) | -91.0 | 4.4 | -75.5 | 4.4 |
| Basic general track minus Academic track (Gymnasium) | -149.9 | 5.9 | -130.1 | 7.4 |
Special-education, Waldorf and upper-secondary vocational programmes are deliberately left unclassified rather than forced into one of the four tracks, which is why the shares do not sum to the full cohort.
The usual caution applies with unusual force here. Students are not randomly assigned to tracks, prior attainment is not measured, and the sorting happened five years before the test. The gap at 15 is real; how much of it the tracks created rather than inherited is not answerable from PISA.
Can this comparison be trusted? A measurement audit
Test scores are placed on a common scale by design. Questionnaire indices are not. Comparing an index across groups assumes the items mean the same thing in every group being compared, and that assumption is testable. It is usually not tested.
Germany’s post-PISA debate leans heavily on questionnaire indices, and mathematics anxiety in particular has been used to explain track differences. If the anxiety items do not function the same way for a Gymnasium student and a basic-track student, then comparing their anxiety means is a category error dressed as a finding. This is testable, so it is tested.
| Model | Chi-square | df | CFI | RMSEA | SRMR | Decision |
|---|---|---|---|---|---|---|
| Configural | 947.50 | 36 | 0.881 | 0.230 | 0.056 | reference |
| Metric | 1092.02 | 51 | 0.879 | 0.195 | 0.059 | supported |
| Scalar | 1224.17 | 66 | 0.874 | 0.174 | 0.062 | supported |
| Strict | 1244.53 | 84 | 0.872 | 0.156 | 0.062 | supported |
Measurement invariance of the mathematics anxiety block across school track. N = 4,607. Estimator MLR. Decision rule after Chen (2007).
Verdict. Strict invariance held; loadings, intercepts, and residual variances are equivalent across groups, supporting comparison of observed means and (co)variances. The categorical re-run reached strict invariance.
With the verdict in hand the comparison can be read. Anxiety is lowest in the Gymnasium and higher in every other track, and the ordering follows attainment closely enough to raise the obvious question of direction, which cross-sectional data cannot settle.
What follows
Three statements are supported and narrow enough to defend.
- The immigrant gap is not mainly a tracking artefact. Roughly a third is socio-economic composition, roughly an eighth is track placement, and more than half is neither.
- The track gap is the largest structure in German 15-year-old attainment, at 130.1 points between the basic general track and the Gymnasium net of socio-economic status, and it applies to a cohort sorted at about age ten.
- Germany’s socio-economic gradient is the steepest in this eight-country series. A policy that reduced the gradient to the OECD average would move more points than any other single change visible in these data.
The next analysis this points to is the within-track immigrant gap by language background and by years since arrival, which the public file supports and which would separate a language effect from a longer-run integration effect. That is a deep-dive rather than a note.
What this note does not claim
- No trend statement. Comparing 2022 with an earlier cycle requires the published link error for that cycle pair to be carried in the variance. No such comparison is made here.
- No causal claim. Every difference is an association measured at one point in time. “Net of ESCS” means one measured index is held constant, not that other things are equal.
- The invariance cascade is fitted without the replicate-weight design, which is standard practice and is stated rather than left implicit.
- Track is reconstructed from the national study-programme code rather than reported directly. The published labels are unambiguous about the academic, comprehensive, intermediate and basic distinctions, but any grouping of programme codes is an interpretation and a different grouping would move the numbers somewhat.
- Germany’s regional structure is invisible here. The public file suppresses the stratum variable for Germany, so no Land-level analysis is possible from this source, and none is attempted.
Method and reproducibility
Data: OECD PISA 2022 Public Use Files, downloaded from webfs.oecd.org on 6 August 2026. Point estimates pooled over ten plausible values by Rubin’s rules; sampling variance from 80 Fay-adjusted replicate weights with a Fay factor of 0.5, computed per plausible value and averaged; imputation variance inflated by (1 + 1/M). Invariance tested as a staged configural, metric, scalar and strict cascade in lavaan with full-information estimation for the rotated questionnaire design, plus a categorical sensitivity re-run on pairwise-present data.
Subgroups are derived from variables the OECD codes identically in every participating system, so this note was produced without country-specific recoding. Where a country’s own sampling strata carry a more policy-relevant structure, that structure is used instead and the note says so.
Enquiries about reproducing the analysis are welcome at milos@centerforpsychology.me. Commissioning the equivalent for another country is consulting work, handled by AdriaMont Consulting DOO at milos@adriamont.me.
All country notes in this series
About
Dr Milos Kankaras is a psychometrician and policy analyst with more than twenty years of international large-scale assessment work, including with the OECD, UNESCO and Eurofound. His published specialism is measurement equivalence and cross-cultural comparability. He holds a PhD in social sciences from Tilburg University.
This note is one of a series of country notes produced from the same analysis code. Published by the Center for Psychology, Podgorica. The Center for Psychology and the AdriaMont Institute are both operated by AdriaMont Consulting DOO, Cetinjski put 36, 81100 Podgorica, Montenegro. Commissioning enquiries go to milos@adriamont.me.
Dr Milos Kankaras | milos@centerforpsychology.me | miloskankaras.com | centerforpsychology.me | ORCID | LinkedIn