This project estimates, for every country it tracks and for the world as a whole, the share of the resident population that is predominantly of European ancestry, each year from 1950 to 2100. This page explains what that means, how the estimates are built, and the rulings that apply across all countries.

What we measure

European ancestry here means ancestry traceable to the indigenous peoples of Europe — the populations present on the continent before the modern era of intercontinental migration.

Genetically, modern Europeans descend from a three-way admixture of Western Hunter-Gatherers, Early European Farmers (originating in Anatolia), and Western Steppe Herders (from the Pontic–Caspian steppe), a model established by ancient-DNA research (Lazaridis et al., 2014; Haak et al., 2015). It is this shared continental ancestry that the metric tracks. The choice is not arbitrary: population-genetic studies consistently resolve humanity into a small number of major continental ancestry clusters — Sub-Saharan African, European, Middle Eastern and North African, South Asian, East Asian, Oceanian, and Native American — each distinct from the others (Li et al., 2008). Just as there are populations indigenous to Oceania, to East Asia, or to the Americas, there is a population indigenous to Europe, and tracking its size and share is a matter of the same demographic interest.

The measure is about ancestry — not nationality, language, religion, or self-identified race. A French citizen of Algerian descent is not counted as European-descended; a Polish citizen living in Britain is. The question is always the same: what share of the people living in a country are predominantly of European indigenous ancestry?

In principle the measure is the share of residents whose own ancestry is predominantly European — a working threshold of at least roughly three-quarters. This number is never settled person by person; it is approximated, group by group, with the best available proxy. For most populations the proxy — ethnicity, national origin, census race — is taken to satisfy the threshold without being tested, and no one’s ancestry is actually measured: the Polish resident just mentioned is counted as European because for such groups the assumption is safe. The threshold becomes an explicit question only for certain populations where it is genuinely in doubt—for example, mixed-ancestry populations in Latin America, where genetic-admixture studies may be needed because census self-identification does not reliably indicate how many people are predominantly European-descended.

Why European-descended rather than “White”?

In everyday language, White and European-descended overlap substantially, but they are not interchangeable. White is a social and administrative category whose boundaries vary between countries, censuses, and historical periods. European-descended gives the project a consistent cross-country standard: ancestry predominantly traceable to the indigenous populations of Europe.

The United States illustrates the problem. Through the 2020 census, federal standards classified Middle Eastern and North African populations as White. The Census Bureau’s geographically based MENA classification includes Arabic-speaking groups such as Egyptians and Jordanians, non-Arabic-speaking groups such as Iranians and Israelis, and ethnic and transnational groups such as Assyrians and Kurds — populations whose origins are outside indigenous European ancestry (U.S. Census Bureau, 2023). Conversely, the modern non-Hispanic White category excludes every Hispanic respondent, including the minority whose ancestry is predominantly European. A White census count can therefore both include people outside our definition and exclude people who meet it.

Self-identification creates a similar mismatch. Among U.S. Jews, 92% described themselves as White and non-Hispanic in Pew’s 2020 survey. Yet genetic studies estimate Ashkenazi Jewish ancestry at roughly half European and half Middle Eastern — below the project’s predominantly European threshold (Pew Research Center, 2021; Carmi et al., 2014; Xue et al., 2017). In parts of South America, people with substantial Indigenous or African admixture may likewise identify as White because racial identity reflects appearance, social status, and local convention as well as ancestry (Paschetta et al., 2021; Ruiz-Linares et al., 2014).

For these reasons, White census categories remain valuable proxies, but they are not treated as the metric itself. The project adjusts them toward a single ancestry definition so that figures from different countries and periods answer the same question.

Why we build our own estimate

No country directly tracks the share of its population that is of European descent. Statistical offices collect things that are close but not identical — census race categories, ethnicity, country of birth, citizenship, migration background. Each is a useful stand-in, but never a perfect match for ancestry.

Official statistics are therefore an input to this project, not its output. They are built for administration, not for this question, and they miss and misclassify people in predictable ways: bundling Middle Eastern and North African populations into “White,” absorbing later-generation immigrants into “native” categories, or recording European immigrants as merely “foreign.” Each country estimate starts from the best available measure and corrects it toward the ancestry definition, documenting every adjustment and its direction.

How the estimate is built

Each figure is a fraction: the number of European-descended people divided by the total number of people living in the country.

The numerator is corrected for who counts as European-descended — adding European-descended people the source leaves out (for example, European immigrants logged only as “foreign background”) and removing people who are not European-descended but the source includes (for example, MENA populations recorded as “White”).

The denominator is the de facto population: everyone physically resident, whether or not officials counted them. This matters because the people most often missed by official counts — undocumented migrants, unregistered refugees, transient workers — are disproportionately not European-descended, so counting only the registered population would overstate the European-descended share.

Building each country’s series

Each country is estimated as a yearly series from 1950 to 2100, but it is built from a smaller set of anchor years — the years for which a real measurement exists. These are usually census years, supplemented by population-register snapshots, large official surveys, and clearly dated demographic breaks such as a refugee surge or a change in citizenship law, together with a present-day estimate and a few projection points. We anchor only to years a country actually has data for, rather than imposing a tidy decade grid, and the values between anchors are filled by interpolation — never entered as if they were measurements.

The proxy used at each anchor is the best official category available at that time, and it changes over the life of a country’s record. The United States, for instance, is read from the total census “White” count before 1980 and from non-Hispanic White alone afterwards, while the United Kingdom has no usable ethnic-group census before 1991 and relies on reasoned reconstruction for the earlier decades. Because the underlying measure shifts, each country’s record is a chain of era-specific proxies rather than one continuous instrument, and every such break is documented.

At each anchor the raw proxy is corrected toward the ancestry definition as described above — subtracting groups recorded as European that fall outside it, adding back European-descended people filed elsewhere, and adjusting the denominator for residents the count missed. Decades with no usable proxy are carried as reasoned back-casts, and years beyond the present are projection scenarios rather than measurements. Each country’s report sets out the proxies used in each era, the corrections applied to them, and the sources behind them.

Example: the United States

Every country’s report carries a full table of corrected anchor values — the auditable arithmetic behind the estimate. Each row begins with the raw proxy, applies every correction with its direction and approximate size, and arrives at a single value. The full table is public: every value can be traced back to its source proxy, and every correction is visible in the arithmetic.

Year Value (%) Value audit trail
1950 84.4 ~89.5% census total White proxy − ~1.8pp Hispanic-white subtraction − ~0.2pp minor MENA / non-European White residual − ~3.2pp Jewish-population correction + ~0.4pp ≥75%-Hispanic add-back − ~0.3pp historical differential-coverage correction = 84.4%.
2000 67.3 69.1% non-Hispanic White alone proxy − ~0.5pp MENA subtraction − ~0.9pp coverage / de facto correction − ~1.9pp Jewish-population correction + ~1.5pp ≥75%-Hispanic add-back = 67.3%.
2020 55.9 57.8% non-Hispanic White alone proxy − ~0.7pp MENA subtraction − ~1.5pp Census coverage / de facto correction − ~1.7pp Jewish-population correction + ~2.0pp ≥75%-Hispanic add-back = 55.9%.
2050 42.5 43.5% Census projection proxy after MENA / de facto corrections − ~1.5pp projected Jewish-population correction + ~1.2pp decaying ≥75%-Hispanic add-back − ~0.7pp fuller de facto / momentum correction = 42.5%.

For the United States, the proxy switches from census total “White” (which includes Hispanic-white and MENA respondents) to non-Hispanic White alone at 1980, when the census first tabulated Hispanic origin cleanly. Neither measure perfectly matches European ancestry. The older White category includes Hispanic, MENA, Jewish, and other populations that require further separation. The later non-Hispanic White category removes most Hispanic residents, including some whose ancestry is predominantly European. The figures are therefore adjusted using local information about ancestry, Hispanic origin, Jewish communities, census coverage, and populations likely to have been missed.

Who counts as European-descended

Most cases are unambiguous and are settled by group membership: the task is to correct each country’s official categories — removing groups recorded as European that fall outside the definition, and adding back those that belong but are recorded elsewhere — not to estimate any individual’s ancestry. Direct appeal to the genetic standard is reserved for populations that are genuinely admixed and whose official classification can at times be misleading, or where the available categories cannot separate people by ancestry at all; for those, the ancestry threshold described under mixed ancestry below decides the question. The rulings that follow cover these harder cases and apply across every country.

Included

  • People of indigenous European ancestry, regardless of citizenship or where they were born.
  • Bosniaks, Albanians, and other Balkan Muslims — the boundary is ancestry, not religion, and these populations are genetically indigenous to Europe.
  • European ethnic minorities living outside their titular nation, such as ethnic Russians in the Baltic states.

Excluded

  • Middle Eastern and North African (MENA) populations, whose ancestry falls outside the European profile.
  • Jewish populations (see mixed ancestry below).
  • Deeply admixed Latin American populations, such as Brazilian pardos and Mexican mestizos (see mixed ancestry below).
  • Turkic populations, such as Turks and the Gagauz, whose origins lie in Central and West Asia — regardless of long residence, citizenship, or local birth.
  • Roma, whose ancestry traces to South Asia, are not treated as European-descended: they remain part of the total resident population but are kept out of the European-descended count, and documented separately rather than silently dropped.

These exclusions hold regardless of citizenship or country of birth: a third-generation citizen of non-European ancestry is counted by ancestry, not by passport.

Mixed ancestry

Mixed ancestry is handled by counting whole people, not fractions of ancestry; no individual is assigned partial credit. What we are after is the share of individuals whose own ancestry clears the roughly-75% threshold. For a group that sits consistently well below that line — Brazilian pardos, Mexican mestizos — genetic-admixture studies settle the matter at the group level: essentially no one clears the threshold, so the whole group is excluded from the European-descended figure, even where it is officially recorded as white, while remaining part of the total resident population against which that figure is measured. For a group consistently well above the line, essentially everyone clears it and the group is counted in. Where a population is internally mixed — a substantial part clearly above the threshold and a substantial part clearly below, as in some Latin American countries — its overall average is not a reliable guide, and the estimate is built instead from the share of members who individually clear the threshold, using regional or origin-group evidence rather than a single national mean. A population’s average ancestry is never itself reported as its European-descended share; doing so would credit fractions of people, which the whole-person rule forbids.

This matters most in Latin America, where self-identification diverges sharply from measured ancestry. The historical and social dynamics of blanqueamiento, or “whitening” — in which higher socioeconomic status and education raise the likelihood that a person reports as white — mean that census figures substantially overstate the share of the population that is genetically predominantly European (Paschetta et al., 2021; Ruiz-Linares et al., 2014). Deeply mixed populations such as Brazilian pardos and Mexican mestizos therefore fall outside the total, even where they are recorded as white.

The same threshold places Jewish populations outside the European-descended total. Ashkenazi Jews carry roughly 50% European ancestry (Carmi et al., 2014; Xue et al., 2017); Sephardi Jews somewhat lower; Mizrahi and Yemenite Jews under 15%; and Ethiopian Jews (Beta Israel) essentially none. Most groups share a Levantine and Middle Eastern founding ancestry that places them outside the European cluster; Ethiopian Jews (Beta Israel) are genetically an East African population with minimal Eurasian admixture of any kind. No Jewish subgroup meets the 75% threshold. Settler-origin populations in the Americas and Oceania, by contrast, sit above it and are treated as mostly, though not entirely, European-descended.

Historical data and projections

For most European countries the series begins near 98–99% in 1950, before large-scale non-European immigration. The decline that follows is tied to concrete events — post-war guest-worker recruitment, decolonisation, changes in immigration law, refugee and asylum flows, EU free movement, and the decisions of national governments to open migration. Each country’s report names the specific causes that bent its own curve. Figures from 1950 to the present rest on censuses, registers, surveys, and genetic studies, corrected as described above for each tracked country.

Figures beyond the present are projections, not measurements. They are reasoned from cohort replacement — older, more European-descended generations being succeeded by younger, less European-descended ones — together with fertility differences and the likely composition of future migration. Where a statistical office projects a related composition, we build on it and convert it to the ancestry definition; the United States Census Bureau, for example, projects population by race and Hispanic origin to 2100, and Statistics Canada projects the country’s visible-minority and ethnocultural composition, both of which we translate into a European-descent path. Where no such projection exists, as in much of Eastern Europe, the estimate rests instead on United Nations and national population projections together with reasoning about cohort replacement, fertility differences, emigration, and the likely composition of future migration. In every case the European-descent figure is our own synthesis. As with any long-range demographic projection, the uncertainty widens the further out it reaches; the values around 2100 should be read as scenarios, not forecasts.

The global figure

The worldwide figure is not an average of the country percentages. It is assembled as a global headcount: the European-descended people in the tracked countries, plus those in other countries with significant European-descended populations (such as Brazil, Mexico, and South Africa), plus a bounded estimate for the rest of the world, all divided by the United Nations’ total world population. This keeps the global share grounded in absolute numbers rather than in an unweighted blend of national rates.

Regional figures

Each tracked country belongs to a region — North America, South America, Western Europe, Northern Europe, Southern Europe, Central Europe, Southeastern Europe, Eastern Europe, and Australasia. A region’s European-descended share is a population-weighted aggregate of its member countries, not a simple average of their rates: in effect the region is treated as one large country, with the European-descended people of its members summed and divided by those members’ combined population, so a populous country weighs more than a small one. The regional denominator is therefore the combined population of the member countries.

Regions include only countries the project tracks; the additional countries used to assemble the global figure are not folded in. This matters most in the Americas, where several countries feed the global figure but where the South America region includes only Argentina and Uruguay. The European regions, by contrast, contain almost all of their constituent countries and so track the geographic region closely.

Metropolitan-area figures

A metropolitan area includes the central city and the surrounding communities that are closely connected to it through commuting, employment, and everyday life. “Greater New York,” for example, means the wider New York–Newark–Jersey City metropolitan area, not New York City alone.

Different countries use different official names for these areas. The United States uses Metropolitan Statistical Areas (MSAs), while Canada uses Census Metropolitan Areas (CMAs). We use these official boundaries whenever possible, even when the map shows a shorter, more familiar name.

Keeping the boundaries consistent over time

Metropolitan boundaries change. Counties and municipalities may be added, removed, divided, merged, or renamed. If we simply used the official metro total published in every census year, part of the apparent demographic change could come from the boundary changing rather than from the population itself.

To avoid this, we choose one boundary and keep it as consistent as possible throughout the series. For the United States, the same July 2023 MSA boundaries are used for the historical figures as well as the present-day ones.

Older censuses did not always publish data for those modern boundaries. In those cases, we rebuild the metro from the smaller areas available in the historical census — usually counties, county equivalents, or towns. We also account for places whose boundaries or administrative status changed over time.

The result answers a consistent question: what share of the people living within this metropolitan footprint were predominantly European-descended at each point in time?

Building the figures

Each metro figure begins with the closest available census or statistical count. Where an official table already covers the correct metropolitan area, we use it. Otherwise, we add together the counties, towns, or other local areas that make up the metro.

We then make the same corrections used for the national figures, but with local evidence wherever possible. These include:

  • removing populations counted as White or European by the source but falling outside this project’s ancestry definition;
  • adding European-descended populations that the source placed in another category;
  • accounting for mixed populations using the project’s predominantly-European threshold; and
  • including people who were living in the metro but were likely missed by the census or population register.

These corrections are not spread evenly across every metropolitan area. Different populations are concentrated in different places. The MENA correction, for example, is particularly important in metros such as Detroit, New York, Washington, Los Angeles, and Chicago, and less important in other metro areas of the United States. Corrections for missed residents are generally larger in major immigration centres than in places with fewer recent arrivals.

When local figures are available, we use them directly. When only a national total is known, we use the best available local evidence—such as ancestry, birthplace, immigration, or community data—to estimate how that population was distributed. Where this requires more uncertainty, the figure is identified as a reasoned estimate rather than a direct count.

This work is repeated for every included metro and every census year. The national percentage is used as a check, but it is not copied onto the metros. The combined metropolitan figures must fit plausibly within the national population, while still allowing places such as Miami, New York, Salt Lake City, or Pittsburgh to follow very different demographic paths.

Figures between census years are filled in gradually between the nearest researched years. Figures for 2050 and 2100 are projections rather than measurements. They combine the national demographic outlook with each metro’s own history and population pattern. As with the country projections, uncertainty grows the further into the future the series extends.

How we use AI

This project uses frontier AI as a research system, not as a primary source of facts. AI coding agents help retrieve, organize, analyze, and test far more material than a single researcher could reasonably process by hand. The underlying evidence still comes from official censuses, statistical databases, academic research, and the other sources cited throughout the project.

For recent decades, the agents make thousands of API calls to official statistical databases, download extensive tables, write code to combine and analyze the records, reconstruct consistent geographic areas, and check the results against published totals. For older periods, much of the evidence survives only in scanned copies of printed reports. AI helps locate these documents, run optical character recognition (OCR), compare the extracted text with the original page images, correct extraction errors, reconstruct damaged tables, and preserve the recovered data for later re-examination.

Models used in the project include Opus 4.8, Fable 5, GPT-5.5, and GPT-5.6 Sol. As new generations become available, they are used to audit earlier work and search for weaknesses that previous models may have missed. The figures remain traceable to their sources, with assumptions and calculations recorded for repeated review. This allows, for example, Fable 5 to independently audit work performed by GPT-5.6 Sol and return any problems for correction. Throughout the process, human judgment sets the methodology and resolves disputed questions.

The purpose of AI is to provide a tireless research and review layer, allowing each new generation of frontier agents to revisit the recorded work and help make the estimates more complete, transparent, and robust over time.

Why we track this

The share of the world’s population that is of European descent is falling rapidly, and in many individual countries a long-standing European-descended majority is receding within a single lifetime. This project treats that as a significant and consequential transformation — one worth measuring carefully and stating plainly, without exaggeration, without understatement.

References

Carmi, S., et al. (2014). Sequencing an Ashkenazi reference panel supports population-targeted personal genomics and illuminates Jewish and European origins. Nature Communications, 5, 4835.

Haak, W., et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207–211.

Lazaridis, I., et al. (2014). Ancient human genomes suggest three ancestral populations for present-day Europeans. Nature, 513, 409–413.

Li, J. Z., et al. (2008). Worldwide human relationships inferred from genome-wide patterns of variation. Science, 319(5866), 1100–1104.

Paschetta, C., et al. (2021). The impact of socioeconomic and phenotypic traits on self-perception of ethnicity in Latin America. Scientific Reports, 11, 12617.

Pew Research Center. (2021). Race, ethnicity, heritage and immigration among U.S. Jews. Jewish Americans in 2020.

Ruiz-Linares, A., et al. (2014). Admixture in Latin America: geographic structure, phenotypic diversity and self-perception of ancestry based on 7,342 individuals. PLOS Genetics, 10(9), e1004572.

U.S. Census Bureau. (2023). 3.5 million reported Middle Eastern and North African descent in 2020.

Xue, J., et al. (2017). The time and place of European admixture in Ashkenazi Jewish history. PLOS Genetics, 13(4), e1006644.