# Does AI Have an IQ Over 130? What AI IQ Scores Mean and Where They Break Down

Source: https://iqcat.vercel.app/en/blog/ai-iq-score
Published: 2026-09-04
Publisher: NekoIQ (https://iqcat.vercel.app)
Tags: IQ基礎, 心理学

> Headlines about an AI scoring an IQ of 136 come from giving a human Mensa-style test to a model. IQ is a relative score based on a human age cohort, and public tests can appear in training data, so AI scores do not mean what human IQ means. Here is the research and its limits.

## Key takeaways

- AI IQ figures in the news come from feeding public human IQ tests such as Mensa Norway to a model and reading the result off a human conversion table.
- IQ expresses a person's position within a human age cohort. An AI has neither an age nor a reference population, so its number does not carry the same meaning as a human IQ.
- Public test items can appear in training data. Scores reportedly fall on a test that has never been posted online.
- In a peer-reviewed study, GPT-3 matched or exceeded humans on a text-based matrix task modeled on Raven's matrices, but that is a result on one task type, not proof of general intelligence.

Headlines announcing that the latest AI has passed IQ 130 are becoming routine. Taken at face value, the number corresponds to the top 2% of people. But what exactly was measured, and how? This article looks at how AI IQ figures are produced and asks, starting from the definition of IQ, whether they can be read the way human scores are.

## How an AI IQ is measured

Most AI IQ figures quoted in the media come from TrackingAI, an independent benchmarking site. It gives models two tests: the public figure-pattern IQ test from Mensa Norway (35 items, 25-minute limit) and an offline test written by a Mensa member that has never been posted on the internet. Raw results are converted to IQ with tables built for humans.

Items are generally verbalized, that is, described in text, and models that accept images are also shown the figures themselves. The site publishes procedural details such as re-asking an item up to ten times if the model refuses to answer.

By this method, 2025 reports put [OpenAI's o3 at the equivalent of IQ 136 on Mensa Norway and Gemini 2.5 Pro at 128](https://ledge.ai/articles/tracking_ai_mensa_iq_test), among others. The figures change with every model update, so any specific value is a snapshot rather than a fixed property.

## IQ is a position within a human population

Return to the definition. IQ is not an absolute quantity of intelligence but a relative score: raw scores from a same-age human population are standardized to mean 100 and SD 15 (see [What is the average IQ?](https://iqcat.vercel.app/en/blog/iq-average)). [IQ 130](https://iqcat.vercel.app/en/iq/130) means roughly the top 2% only because that reference population exists.

An AI has no age and no cohort of same-age AIs to be compared with. Reading its raw score off a human table is a translation that says a human with this many correct answers would score 130. It does not mean the AI possesses the intelligence of the top 2% of humans. The same logic explains why dog and cat intelligence cannot be expressed on the human IQ scale (see [Where do cats rank in animal intelligence?](https://iqcat.vercel.app/en/blog/animal-iq-ranking)).

## Public tests can be in the training data

A second limit is test security. Anyone can take the Mensa Norway test online, so the items and answers may be part of a model's training data. TrackingAI keeps a separate never-published test precisely to avoid this contamination.

In human testing, knowing the items in advance raises the score without raising intelligence (see [Does practicing IQ tests help?](https://iqcat.vercel.app/en/blog/iq-test-practice)). For AI, the training corpus is so vast that distinguishing solving a new problem from reproducing a seen one becomes fundamentally hard.

## What peer-reviewed research has shown

Peer-reviewed work exists. [Webb and colleagues (2023)](https://doi.org/10.1038/s41562-023-01659-w) directly compared GPT-3 (text-davinci-003) with human participants on tasks including a text-based matrix reasoning problem built on the rule structure of Raven's Standard Progressive Matrices, as well as letter-string analogies. GPT-3 matched or exceeded humans in most settings, which the authors describe as an emergent capacity for abstract pattern induction.

What the study showed, however, is performance on specific task formats. The paper itself does not claim to have measured intelligence as a whole. Human IQ tests are valid because their scores correlate consistently with external criteria such as academic and occupational outcomes (see [IQ and school grades](https://iqcat.vercel.app/en/blog/iq-and-grades)); AI scores have no such backing yet.

## Same items, different solving process

Raven-type items ask you to spot the rule in a grid of figures, such as rotation, counting, or overlay, and fill the blank (see [What are Raven's Progressive Matrices?](https://iqcat.vercel.app/en/blog/ravens-progressive-matrices)). Humans solve them with visual working memory and reasoning; a model receiving a verbalized item completes statistical patterns over sequences of symbols.

Because the underlying processing differs even when the number of correct answers is the same, the IQ figure supports neither the conclusion that AI is smarter than people nor that it is dumber. What can be compared is only the accuracy rate on this task format.

## Why take an IQ test in the age of AI

The purpose of an IQ test does not change because AI can solve figure puzzles. The goal is to learn where your reasoning stands among people, which is independent of AI performance. If anything, as routine work moves to AI, knowing which kinds of reasoning you are good at becomes more valuable. How leaning on AI affects thinking is covered in [Does relying on AI make you dumber?](https://iqcat.vercel.app/en/blog/ai-cognitive-debt).

The [free NekoIQ IQ test](https://iqcat.vercel.app/en/start) scores 20 questions across spatial reasoning, pattern recognition, logical reasoning, and classification (about 10 minutes) with item response theory and reports an estimated IQ, deviation score, and top percentage on the mean-100, SD-15 scale (see [What is IRT?](https://iqcat.vercel.app/en/blog/what-is-irt)). The numbers are estimates, not the results of a formal psychological assessment.

## Frequently asked questions

### What is the IQ of AI?

According to the independent site TrackingAI, OpenAI's o3 was reported in 2025 to score the equivalent of IQ 136 on the public Mensa Norway test and Gemini 2.5 Pro about 128. These are reference values read off a human conversion table and do not mean the same thing as a human IQ.

### Does an AI IQ of 130 mean the top 2%, as it would for a person?

No. IQ 130 means the top 2% only when scores are standardized against a human age cohort. An AI has no age and no reference population, and it may have seen the public test items during training, so the scales are not comparable.

### Is there still a point in taking an IQ test that AI can pass?

Yes. The purpose of an IQ test is to learn where you stand among other people, which is independent of whether an AI can solve the items. As routine work shifts to AI, knowing your own reasoning strengths becomes more valuable, not less.

## Disclaimer

NekoIQ provides estimates for entertainment and self-understanding. It is not a medical, clinical, or official psychological assessment. Percentages are derived from a normal distribution with mean 100 and SD 15.
