Reasoning benchmark

# Adversarial NLI leaderboard

> Adversarial NLI results for 9 AI models, led by GPT-3.5-turbo at 58.1%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/anli
- Last updated: 2026-10-10
- Title: Adversarial NLI Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-3.5-turbo has the highest published Adversarial NLI score on Noometry at 58.1%, out of 9 models with results.

Last verified October 10, 2026

## About Adversarial NLI

Natural-language inference examples collected adversarially against earlier models.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2019
- **Format:** Classification
- **Unit:** Percent (random guessing ≈ 33.3%)
- **Official site:** [github.com](https://github.com/facebookresearch/anli)

## Top 9 models

Top models on Adversarial NLI

1.  GPT-3.5-turbo 58.1%
2.  Phi 3 Small 8k Instruct 58.1%
3.  Llama 3-8B 57.3%
4.  phi-3-medium 14B 55.8%
5.  Mixtral 8x7B 55.2%
6.  Phi 3 Mini 4k Instruct 52.8%
7.  Gemma 7B 48.7%
8.  Mistral 7B 47.1%
9.  Phi-2 42.5%
10.  354045505560

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

Adversarial NLI results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | [OpenAI](https://noometry.com/providers/openai) | 58.1% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [Phi 3 Small 8k Instruct](https://noometry.com/models/phi-3-small-8k-instruct) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 58.1% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [Llama 3-8B](https://noometry.com/models/llama-3-8b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 57.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [phi-3-medium 14B](https://noometry.com/models/phi-3-medium-14b) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 55.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Mixtral 8x7B](https://noometry.com/models/mixtral-8x7b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 55.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [Phi 3 Mini 4k Instruct](https://noometry.com/models/phi-3-mini-4k-instruct) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 52.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [Gemma 7B](https://noometry.com/models/gemma-7b) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 48.7% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [Mistral 7B](https://noometry.com/models/mistral-7b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 47.1% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Phi-2](https://noometry.com/models/phi-2) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 42.5% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [GPT-3.5-turbo vs Phi 3 Small 8k Instruct](https://noometry.com/compare/gpt-3-5-turbo-vs-phi-3-small-8k-instruct)
-   [GPT-3.5-turbo vs Llama 3-8B](https://noometry.com/compare/gpt-3-5-turbo-vs-llama-3-8b)
-   [GPT-3.5-turbo vs phi-3-medium 14B](https://noometry.com/compare/gpt-3-5-turbo-vs-phi-3-medium-14b)
-   [GPT-3.5-turbo vs Mixtral 8x7B](https://noometry.com/compare/gpt-3-5-turbo-vs-mixtral-8x7b)
-   [Phi 3 Small 8k Instruct vs Llama 3-8B](https://noometry.com/compare/llama-3-8b-vs-phi-3-small-8k-instruct)
-   [Phi 3 Small 8k Instruct vs phi-3-medium 14B](https://noometry.com/compare/phi-3-medium-14b-vs-phi-3-small-8k-instruct)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [EBR-Bench](https://noometry.com/benchmarks/ebr-bench)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)

## Frequently asked questions

### What does Adversarial NLI measure?

Natural-language inference examples collected adversarially against earlier models.

### Which model has the highest Adversarial NLI score?

As of October 2026, GPT-3.5-turbo has the highest published Adversarial NLI score on Noometry at 58.1%, out of 9 models with results.

### What is the best open-weight model on Adversarial NLI?

Phi 3 Small 8k Instruct has the highest Adversarial NLI accuracy among open-weight models at 58.1%, ranking 2 of 9 overall.

### Cite this page

Noometry. (2026). Adversarial NLI leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/anli

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/anli.md).
