Uzbekistan has the flattest socio-economic gradient in PISA, and a language gap larger than it
In a first-time participant that the OECD’s own figures make look unusually equitable, the language a child is taught in predicts their result better than their family background does.
Summary
Uzbekistan took part in PISA for the first time in 2022, and its published equity figures are striking. Socio-economic status accounts for 2% of the variation in mathematics performance, against 15% on average across OECD countries, and the gap between the top and bottom socio-economic quarters is 22 score points where the OECD average is 93. On the OECD’s own headline equity indicators, Uzbekistan looks like one of the most equitable systems ever measured.
This note takes that result seriously by asking what else is in the data, and finds something the equity headline hides.
What the data shows.
- Students assessed in Russian scored 30.2 points above students assessed in Uzbek (6.9). That single gap is larger than the entire distance between the top and bottom socio-economic quarters.
- It is not composition. Net of socio-economic status the Russian-medium advantage is 25.6 points, and net of socio-economic status and school location together it is 21.0 points (8.4).
- Karakalpak-medium students, 1.9% of the cohort, sit at -5.2 points from the Uzbek-medium majority (12.4), a difference this sample cannot distinguish from zero.
- The socio-economic gradient is genuinely flat: 9.5 score points per unit of the PISA index (1.2), against figures three to four times larger in the European systems in this series.
There is a second reading of the flat gradient that a country note owes its reader. A socio-economic index built from parental occupation, education and household possessions discriminates poorly when most of a population sits in a narrow band near the bottom of the international scale, and Uzbekistan’s mean on that index is far below the OECD average. A flat gradient can mean an equitable system; it can also mean an index with little room to move. Distinguishing the two is a measurement question, not a policy one, and it should be settled before the equity result is used to argue anything.
Why this is the question for Uzbekistan
A first cycle is the moment when a country’s PISA narrative is set, and it is usually set from two or three headline numbers. Uzbekistan’s headline numbers are a low mean and remarkable equity, and that combination invites a specific and comfortable story about a system that serves all its children equally, if not yet well.
The public microdata supports a more specific and more useful story. Uzbekistan assesses in three languages, and the language a school teaches in is associated with a difference larger than the one the equity headline is built on. That is not visible in any published indicator, because no published indicator breaks a country down by language of assessment.
It is also the kind of finding that a first-time participant most needs, because it identifies a structural feature to act on rather than a rank to react to.
The data, and how it was handled
7,293 students in 202 schools, representing a weighted population of 482,059 15-year-olds, from the OECD’s PISA 2022 student and school Public Use Files. National mean in mathematics: 363.9 points (SE 2.0).
Achievement is imputed rather than measured, so estimates are pooled by Rubin’s rules over all ten plausible values. The sample is clustered in schools, so standard errors come from the 80 Fay-adjusted replicate weights the OECD ships for the purpose. Omitting either understates the uncertainty, usually by a factor of two or more.
Before presenting anything new, this note reproduces figures the OECD has already published for this country.
| Published by the OECD | Published | This analysis | SE |
|---|---|---|---|
| Advantaged minus disadvantaged, mathematics | 22 | 22.38 | 3.43 |
| Share of mathematics variance accounted for by ESCS | 0.02 | 2.0% | 0.005 |
| Girls minus boys, reading | 22 | 21.71 | 1.82 |
Source: OECD, PISA 2022 Results (Volume I and II) Country Note: Uzbekistan, published 5 December 2023, https://www.oecd.org/en/publications/pisa-2022-results-volume-i-and-ii-country-notes_ed6fbcc5-en/uzbekistan_2bb94bf1-en.html (read 6 August 2026)
Three languages of assessment, one of which is different
6,584 students were assessed in Uzbek, 565 in Russian and 114 in Karakalpak.
| Language of assessment | Students | Share | Mean, mathematics | SE | Mean ESCS |
|---|---|---|---|---|---|
| Uzbek | 6,584 | 90.2% | 361.7 | 2.1 | -0.74 |
| Russian | 565 | 7.6% | 391.9 | 6.7 | -0.10 |
| Karakalpak | 114 | 1.9% | 356.5 | 12.4 | -0.75 |
The socio-economic difference between the Uzbek-medium and Russian-medium groups is real, which is why the adjustment matters and why it is reported. It does not account for the gap.
| Contrast | Unadjusted | SE | Net of ESCS | SE |
|---|---|---|---|---|
| Russian minus Uzbek | 30.2 | 6.9 | 25.6 | 7.1 |
| Karakalpak minus Uzbek | -5.2 | 12.4 | -5.1 | 12.7 |
Holding school location constant as well is the second obvious check, since Russian-medium schooling is more urban. That leaves 21.0 points (8.4). Comparing a Russian-medium and an Uzbek-medium student of the same socio-economic background in the same kind of settlement, the gap is smaller than the headline and still substantial.
What produces it is not identified by these data, and the plausible candidates differ enormously in what they would imply: teacher supply and qualification, textbook and curriculum quality in each language, selection of families into Russian-medium schools on characteristics the socio-economic index does not capture, or the translation and adaptation of the assessment itself. That last possibility is the one a psychometrician is obliged to name, and it is testable.
The rural majority
Nearly two-thirds of Uzbekistan’s 15-year-olds are in village or rural schools, which is the highest share in this eight-country series and makes the rural figure the national figure rather than a subgroup.
| School location | Students | Share | Mean, mathematics | SE | Mean ESCS |
|---|---|---|---|---|---|
| Village or rural | 4,660 | 64.1% | 358.5 | 2.6 | -0.86 |
| Town | 773 | 10.0% | 368.5 | 7.5 | -0.64 |
| City | 1,860 | 25.9% | 375.7 | 4.0 | -0.28 |
| Contrast | Unadjusted | SE | Net of ESCS | SE |
|---|---|---|---|---|
| Village or rural minus City | -17.2 | 4.7 | -12.8 | 4.8 |
| Town minus City | -7.2 | 8.4 | -4.6 | 8.1 |
The rural and urban gap is smaller than the language gap, and smaller than the equivalent gap in most of the other systems in this series. In a country where the rural population is the majority, that is a genuinely good result and is worth separating from the flat-gradient finding, because it does not depend on the socio-economic index behaving well.
Can this comparison be trusted? A measurement audit
Test scores are placed on a common scale by design. Questionnaire indices are not. Comparing an index across groups assumes the items mean the same thing in every group being compared, and that assumption is testable. It is usually not tested.
The language finding makes the measurement audit unavoidable. If the questionnaire items function differently in the Uzbek and Russian versions, then a language comparison of any index built from them is comparing translations rather than students, and the same doubt attaches by extension to the achievement instrument.
| Model | Chi-square | df | CFI | RMSEA | SRMR | Decision |
|---|---|---|---|---|---|---|
| Configural | 1671.88 | 27 | 0.717 | 0.227 | 0.124 | reference |
| Metric | 1711.86 | 37 | 0.705 | 0.198 | 0.131 | rejected |
| Partial-metric | 1709.88 | 35 | 0.711 | 0.202 | 0.129 | partial |
| Scalar | 1857.13 | 47 | 0.697 | 0.178 | 0.132 | rejected |
| Partial-scalar | 1798.93 | 43 | 0.703 | 0.185 | 0.131 | partial |
| Strict | 1861.67 | 59 | 0.690 | 0.161 | 0.133 | rejected |
| Partial-strict | 1809.04 | 57 | 0.695 | 0.162 | 0.133 | partial |
Measurement invariance of the sense of belonging at school block across language of assessment. N = 6,749. Estimator MLR. Decision rule after Chen (2007).
Verdict. Highest level supported: partial strict. The categorical re-run reached metric invariance.
The two estimators disagree, and the disagreement is the honest headline. The continuous run reaches partial strict invariance; the categorical run, treating the four-point items as ordered, gets no further than metric. When a verdict depends on the estimator, the correct report is the weaker one: loadings can be treated as equivalent across the assessment languages, so relationships involving belonging may be compared, but its means may not.
That is directly relevant to the achievement finding. If the questionnaire instrument does not achieve intercept equivalence across the Uzbek and Russian versions, the possibility that part of the achievement gap is also a translation and adaptation effect is raised rather than dismissed. This analysis cannot settle it and does not pretend to.
Absolute fit is poor in all groups, with a configural CFI near 0.72, so the belonging block is not a clean single factor in Uzbekistan either. A first-cycle participant would be well advised to treat the questionnaire scales as unvalidated locally until a national psychometric study says otherwise.
What follows
Three statements are supported and narrow enough to defend.
- The language of assessment is associated with a larger difference in Uzbekistan than socio-economic status is, and it survives adjustment for socio-economic status and school location.
- The published equity result is real as measured, but a flat gradient in a distribution compressed near the bottom of the international socio-economic scale is not by itself evidence of an equitable system. The two explanations should be separated before the result is used.
- The rural and urban gap is modest by international standards, which in a system where rural schools educate the majority is the most encouraging finding in these data.
For a first-time participant the highest-value follow-up is not another cycle, it is a differential-item-functioning study across the three assessment languages. It would establish whether the language gap is an educational finding or a measurement one, and every subsequent Uzbek PISA analysis depends on the answer.
What this note does not claim
- No trend statement. Comparing 2022 with an earlier cycle requires the published link error for that cycle pair to be carried in the variance. No such comparison is made here.
- No causal claim. Every difference is an association measured at one point in time. “Net of ESCS” means one measured index is held constant, not that other things are equal.
- The invariance cascade is fitted without the replicate-weight design, which is standard practice and is stated rather than left implicit.
- The Karakalpak-medium group is small and its estimates are correspondingly imprecise; it is reported for completeness rather than for interpretation.
- Immigrant background could not be analysed: too few students fall outside the native-born category for a reportable comparison.
- This note is published in the place of a South Asian country note. No South Asian system participated in PISA 2022; the substitution and its reasoning are documented separately in the series coverage note.
Method and reproducibility
Data: OECD PISA 2022 Public Use Files, downloaded from webfs.oecd.org on 6 August 2026. Point estimates pooled over ten plausible values by Rubin’s rules; sampling variance from 80 Fay-adjusted replicate weights with a Fay factor of 0.5, computed per plausible value and averaged; imputation variance inflated by (1 + 1/M). Invariance tested as a staged configural, metric, scalar and strict cascade in lavaan with full-information estimation for the rotated questionnaire design, plus a categorical sensitivity re-run on pairwise-present data.
Subgroups are derived from variables the OECD codes identically in every participating system, so this note was produced without country-specific recoding. Where a country’s own sampling strata carry a more policy-relevant structure, that structure is used instead and the note says so.
Enquiries about reproducing the analysis are welcome at milos@centerforpsychology.me. Commissioning the equivalent for another country is consulting work, handled by AdriaMont Consulting DOO at milos@adriamont.me.
All country notes in this series
About
Dr Milos Kankaras is a psychometrician and policy analyst with more than twenty years of international large-scale assessment work, including with the OECD, UNESCO and Eurofound. His published specialism is measurement equivalence and cross-cultural comparability. He holds a PhD in social sciences from Tilburg University.
This note is one of a series of country notes produced from the same analysis code. Published by the Center for Psychology, Podgorica. The Center for Psychology and the AdriaMont Institute are both operated by AdriaMont Consulting DOO, Cetinjski put 36, 81100 Podgorica, Montenegro. Commissioning enquiries go to milos@adriamont.me.
Dr Milos Kankaras | milos@centerforpsychology.me | miloskankaras.com | centerforpsychology.me | ORCID | LinkedIn