1(function(){try{var e=typeof window<"u"?window:typeof global<"u"?global:typeof globalThis<"u"?globalThis:typeof self<"u"?self:{},t=new e.Error().stack;t&&(e._sentryDebugIds=e._sentryDebugIds||{},e._sentryDebugIds[t]="5a8e3ebd-7734-40ca-9b7a-b1a49575ab7a",e._sentryDebugIdIdentifier="sentry-dbid-5a8e3ebd-7734-40ca-9b7a-b1a49575ab7a")}catch{}})();const n=`# Lesson 6.2: Emergent Behaviour and the Current Model Landscape 2 3**Series:** Series 6 â What LLMs Can and Cannot Do 4**Estimated Duration:** 2 hours 5**Prerequisites:** Lesson 2.5 (early scaling behaviour); Series 3, 4, 5; Lesson 6.1 6 7--- 8 9## Lesson Overview 10 11The previous lesson surveyed what LLMs can and cannot do. This one asks how we got here, and what the frontier looks like. Two threads matter. 12 13The first is **emergence**: certain capabilities appear in larger models without being present in smaller ones, even when both are trained the same way. Wei et al. (Google, 2022) catalogued these in *Emergent Abilities of Large Language Models*, and the framing shaped how the field thinks about **scaling laws**. Schaeffer et al. (2023) argued that what looks like emergence may be a measurement artefact. The debate is live. 14 15The second is **the current model landscape**: which labs ship which models, what distinguishes them, what "frontier" means in 2025, and how to read **benchmark** results without being misled by them. 16 17By the end you should be able to name three emergent capabilities, identify the major frontier models, explain the open-vs-closed-weight debate, and read lab announcements with calibrated scepticism. 18 19--- 20 21## What Emergence Means 22 23Start with the empirical finding. As models grow larger (more parameters, more training data, more compute), their capabilities improve in ways that are mostly continuous: bigger model, slightly better at most things. But for some specific capabilities, the improvement is *not* continuous. The capability is essentially absent in smaller models (performance is no better than chance) and then, at some scale threshold, the capability appears, and the larger models perform substantially above chance. 24 25Wei et al. (2022) catalogued several examples. **Three-digit arithmetic**: small models (under 10 billion parameters) cannot reliably perform three-digit addition; larger models can. **Word unscrambling**: small models cannot solve the task; larger models can. **Logical reasoning on word problems**: small models perform at chance; larger models perform substantially above chance. The pattern across these cases is a sharp **phase transition** at some scale rather than a smooth improvement curve. 26 27The framing the paper popularised: these capabilities are *emergent*, emerging at scale without being explicitly trained, in a way that suggests scaling produces qualitative rather than just quantitative changes. The framing was striking because it implied that something important and not-fully-understood was happening as models grew, and that the future of capability would not be predictable from extrapolating small-model performance. 28 29Subsequent work has refined this picture. **In-context learning** (the ability to learn from a few examples in the prompt without further training) is the most commercially important **emergent capability**. Brown et al.'s GPT-3 paper (OpenAI, 2020) was the original demonstration: a 175-billion-parameter model could be given a few examples of a task in the prompt and then perform the task on new input
29s. Smaller models could not do this; the capability emerged at scale. 30 31**Chain-of-thought reasoning** (Wei et al., 2022) is another commonly-cited emergent capability. The improvement from chain-of-thought prompting is much larger in big models than in small ones; below some scale, the technique provides no benefit, while above the scale it provides large benefits. 32 33*Multi-step instruction following* (the ability to handle prompts with multiple sub-tasks, or instructions that require maintaining state across turns) is similarly more reliable in larger models than in smaller ones, and the improvement is non-linear. 34 35[MANIM ANIMATION PLACEHOLDER: EMERGENCE CURVE. A 2D graph. X-axis: model scale (parameters, log scale). Y-axis: accuracy on a specific task. Multiple curves are shown, each for a different task. A "smooth" curve rises gradually from low to high accuracy as scale increases, typical for tasks where capability scales smoothly. A "sharp" curve stays flat at chance accuracy for several orders of magnitude, then jumps suddenly to high accuracy, typical for emergent capabilities. The contrast between the curves is the visual point: emergence is the sharp-jump pattern. The takeaway: capability gains are not always smooth, and the jumpy gains are what got the field's attention.] 36 37The figure shows what **emergence** looks like empirically. The sharp-jump pattern is what Wei et al. identified, and the surprising fact that some capabilities show this pattern rather than smooth improvement is what made the framing compelling. 38 39--- 40 41## The Counter-Argument: Are Emergent Abilities a Mirage? 42 43The counter-paper from Schaeffer et al. (2023) made a specific technical argument: the *appearance* of emergence may be an artefact of how the metrics are constructed, not a property of the underlying capability. 44 45The argument: many of the metrics used to measure capability are *discontinuous*. For example, on a math **benchmark**, the metric might be "exact match," meaning the model gets the answer right or wrong with no partial credit. A model that produces an answer that is almost right (one digit off) gets the same score as a model that produces a wildly wrong answer (orders of magnitude off). As model scale grows, the model's outputs are gradually getting more accurate; but until they cross the threshold of "exactly right," the metric reports zero. Once they cross the threshold, the metric jumps to high values. 46 47If the underlying capability is improving smoothly but the metric is discontinuous, the resulting curve will look like **emergence**. Schaeffer et al. demonstrated this empirically by re-evaluating the Wei et al. tasks with continuous metrics (e.g., partial credit, edit distance to the correct answer). With continuous metrics, many of the supposedly emergent capabilities show smooth improvement curves rather than sharp jumps. 48 49This is not a knock-down argument. Some capabilities still show jumps even with continuous metrics; the framing of emergence is still useful for some phenomena. But the paper changed how the field talks about emergence. The current consensus, roughly: some capabilities look more emergent than they actually are because of measurement choices; some capabilities are genuinely emergent in a way that survives careful measurement; the practical question of "what scale of model do I need for capability X" is best answered empirically rather than from theoretical predictions. 50 51The debate is worth understanding because it shapes how to read **scaling law** claims. A paper that reports "capability X emerges at 70B parameters" might mean: 52- The capability genuinely jumps from absent to present at 70B (emergent in the strong sense). 53- The capability is improving smoothly but the metric is discontinuous (Schaeffer-style mirage). 54- The capability is being measured in a way that conflates two different things, and the apparent emergence is the result of the conflation. 55 56A reader who understands the debate can ask which interpretation applies to a given claim, rather than taking the emergence framing at face value. 57 58--- 59 60## A Detailed Walk-Through of In-Context Learning 61 62To make the most important emergent capability concrete, walk through what happens when a model performs **in-context learning** on a specific task. 63 64**The setup.** A user wants to extract structured information from product reviews. They want to take a review and produce a JSON object with fields for product name, rating, sentiment summary, and key complaints. The user has not fine-tuned a model for this task; they will rely on in-context learning. 65 66**The prompt.** The user constructs a prompt with three examples and a new instance: 67 68> *Extract structured information from each review. Output a JSON object with product_name, rating, sentiment, and complaints fields.* 69> 70> *Review: "Great laptop, screen is amazing, but battery dies fast. 4/5"* 71> *Output: {"product_name": "laptop", "rating": 4, "sentiment": "mostly positive", "complaints": ["short battery life"]}* 72> 73> *Review: "Headphones broke after a week. Sound quality is good but build is terrible. 2/5"* 74> *Output: {"product_name": "headphones", "rating": 2, "sentiment": "negative", "complaints": ["broke after a week", "poor build quality"]}* 75> 76> *Review: "Coffee machine works perfectly. Easy to use and clean. 5/5"* 77> *Output: {"product_name": "coffee machine", "rating": 5, "sentiment": "positive", "complaints": []}* 78> 79> *Review: "The phone is fast but the camera is mediocre and it's expensive for what you get. 3/5"* 80> *Output:* 81 82**What the model does.** The model sees the pattern across the three examples: input is a review with rating, output is JSON with specific fields. The model's pattern-completion machinery generalises: given the new review, it produces *{"product_name": "phone", "rating": 3, "sentiment": "mixed", "complaints": ["mediocre camera", "expensive"]}*. 83 84**What happened mechanistically.** The model has not been trained on this specific task. The transformer's attention has identified the pattern from the three examples and applied it to the new instance. The "learning" happens entirely within the forward pass; no parameters are updated. By the time the model sees the *Output:* **token** at the end, it has implicitly inferred the task structure from the examples and is generating tokens conditioned on that inference. 85 86**Why this is striking.** A small model fails at this task. Given the same prompt, a small model might produce JSON-shaped output that is missing fields, has wrong field names, or reverts to producing free text. The capability, *infer task structure from examples and apply it*, appears at scale. 87 88**What can go wrong.** Even at scale, **in-context learning** has failure modes. If the examples are slightly inconsistent (one of them has a different field name than the others), the model may produce inconsistent output. If the new instance is substantially different in shape from the examples, the model may produce output that does not match the examples' format. If the task requires reasoning that the examples did not exemplify, the model may fail on the reasoning step while succeeding on the format. 89 90**The skill of using in-context learning.** Practitioners who use **few-shot prompting** effectively learn to be careful about example consistency, example diversity, example coverage, and example ordering. The choice of examples is not just illustrative; it is the entire training signal the model has for the task. A poorly-chosen set of examples produces unreliable output;
90 a well-chosen set produces reliable output. 91 92This is why prompt engineering is a real discipline. The prompts are the model's training data for the specific task; the careful design of prompts is exactly analogous to the careful design of fine-tuning datasets, just at a smaller scale and faster cadence. 93 94For an annotator: **in-context learning** is what allows production deployments to handle many specialised tasks without dedicated fine-tuning. Understanding why and when ICL works is part of understanding why some applications need fine-tuning and some do not. The deployment team's choice between ICL and fine-tuning is partly a function of the annotation budget (fine-tuning requires preference data, while ICL just requires careful prompting) and partly a function of the task's complexity. 95 96--- 97 98## A Historical Tour of Scaling and Emergence 99 100To understand where the current emergence framing came from, it helps to walk through the historical arc that produced it. The story has several named milestones, each of which shifted the field's understanding of what scaling does. 101 102**The Kaplan scaling laws (OpenAI, 2020).** Kaplan et al.'s *Scaling Laws for Neural Language Models* was the first systematic empirical study of how language model performance scales with compute, data, and parameters. The headline finding: across many orders of magnitude, the loss decreases predictably with scale. The loss-vs-scale curves are smooth power laws, not jumps. This finding shaped how OpenAI and other labs allocated resources, since bigger was almost always better and the gains were predictable. 103 104The Kaplan paper looked at *training loss*, which is one specific way to measure model performance. Loss decreases smoothly. What about specific capabilities? That question was less well-addressed in the original Kaplan work, and it is precisely where the emergence debate later played out. 105 106**GPT-3 and the in-context learning surprise (OpenAI, 2020).** Brown et al.'s GPT-3 paper demonstrated that a 175-billion-**parameter** model could do something smaller models could not: learn new tasks from examples in the prompt without further training. This was a striking finding. **In-context learning** was not what GPT-3 was trained for; it emerged as a side-effect of scale. The GPT-3 paper is also the moment when "**few-shot prompting**" became a recognised technique. 107 108**The Wei et al. emergence paper (Google, 2022).** Wei et al. catalogued several capabilities that show the sharp-jump pattern: three-digit arithmetic, word unscrambling, logical-deduction word problems. The catalogue made "**emergence**" a coherent framing: a class of capabilities with similar scale-dependent behaviour, rather than a single capability emerging in isolation. The paper became the canonical reference for the scaling-produces-qualitative-change view. 109 110**The Chinchilla scaling laws (DeepMind, 2022).** Hoffmann et al.'s *Training Compute-Optimal Large Language Models* (the **Chinchilla** paper) revisited the Kaplan scaling laws and found a different optimal trade-off between data and parameters. The original Kaplan work suggested scaling parameters faster than data; Chinchilla suggested scaling them roughly proportionally. The implication: many large models from 2020â2021 (GPT-3, Megatron, Gopher) were *under-trained on data* relative to their **parameter count**. A smaller, more-data-trained model could match or beat them. 111 112This finding had substantial impact on the field. Subsequent training runs followed **Chinchilla**-style ratios (roughly 20 **tokens** per parameter), and the resulting models, at the same parameter count, were measurably better than their pre-Chinchilla equivalents. The combination of "**compute-optimal training**" and continued parameter scaling has driven much of the capability improvement of 2023â2025. 113 114**The Schaeffer counter-paper (2023).** As discussed above, Schaeffer et al. argued that some apparent emergence is a measurement artefact. The paper's contribution was not to deny that scale matters but to clarify what scale actually produces. Many "emergent" capabilities, when measured continuously, show smooth improvement curves; the sharp jumps are partly an artefact of using exact-match metrics. 115
116**The reasoning-model paradigm (OpenAI, 2024).** The **o1** family introduced a new dimension of scaling: scaling at *inference time* rather than just at training time. The o1 paper showed that giving the model more thinking-time per query produces large gains on hard tasks, even with the same trained weights. This is "**test-time compute**" scaling and it has become a major axis of capability improvement separate from "more training compute." 117 118**The current frame (2025).** The field now thinks about scaling along several dimensions simultaneously: **parameter count** (more weights), training data (more **tokens**), training compute (more flops applied at training), and inference compute (more flops applied per query). Each axis has its own scaling laws and its own returns. Frontier capability is moving along all of them at once. 119 120The historical arc tells you something important: the emergence framing has evolved as the field has learned more. The strong "capability appears at scale thresholds" framing of 2022 has been refined by the 2023 measurement-artefact discussion, and is now being supplemented by the 2024 inference-time scaling story. A current practitioner needs to hold all three in mind: there are genuine **emergent capabilities**, some apparent emergence is measurement-driven, and capability improvement now happens along multiple axes that each have their own scaling behaviour. 121 122For an annotator's purposes: the historical arc helps explain why certain capabilities (in-context learning, chain-of-thought reasoning, multi-step instruction following) became reliable at certain points and not earlier. Understanding the arc is part of understanding why the current frontier looks the way it does. 123 124--- 125 126## In-Context Learning: The Most Important Emergent Capability 127 128Of all the emergent capabilities, **in-context learning (ICL)** is the most consequential commercially. It is what makes prompting a coherent discipline; it is what allows production systems to adapt to new tasks without fine-tuning; it is the mechanism behind **few-shot prompting** and many of the practical capabilities that make modern LLMs useful. 129 130The basic phenomenon: given a prompt that contains several examples of a task, plus a new instance of the task, the model can perform the task on the new instance. The model has not been fine-tuned on this task; the entire "learning" happens within the context window, just from the examples. 131 132A concrete example. The prompt: 133 134> *Translate from English to Spanish.* 135> 136> *English: The cat is on the mat.* 137> *Spanish: El gato está en el tapete.* 138> 139> *English: I want to learn to cook.* 140> *Spanish: Quiero aprender a cocinar.* 141> 142> *English: Where is the train station?* 143> *Spanish:* 144 145The model continues with *"¿Dónde está la estación de tren?"*, a correct translation. The model has not been fine-tuned on translation today; it has been shown the pattern by the in-context examples, and it has generalised. This is **in-context learning** at its most basic. 146 147The capability scales with model size. Small models do not show meaningful ICL; they tend to generate plausible-looking text but fail to apply the pattern. Large models do, and the larger the model, the better the ICL on harder tasks. The transition is sharp enough that it was one of the original **emergent capabilities** Wei et al. catalogued. 148 149Why ICL works is not fully understood. The leading theoretical account, from work by researchers at Anthropic and Stanford, is that the transformer architecture implicitly performs something like Bayesian inference: each example in the prompt updates the model's effective distribution over the task being asked. The math has been worked out for simplified versions of transformers; whether it applies fully to frontier models is contested. The empirical fact (ICL works, scales with model size, and is the mechanism behind **few-shot prompting**) is robust enough that the theoretical question is less important for practical purposes than the engineering question of how to use it well. 150 151--- 152 153## Practice Exercise 6.2.1 154*Allow 15â20 minutes.* 155 156Spend at least fifteen minutes: 157 1581. Pick a frontier chat model and a smaller model (Mistral 7B Instruct or a similar smaller open model, accessed via any API). Design three tasks that probe different kinds of capability: one factual recall task, one in-context learning task (with a few examples in the prompt), and one chain-of-thought reasoning task. 1592. Run all three tasks on both models. Record the responses and your assessment of accuracy. 1603. For each task, note whether the gap between the larger and smaller model is smooth (the smaller is somewhat worse but in the same ballpark) or sharp (the smaller is essentially failing while the larger succeeds). The sharp-gap tasks are candidates for emergent capabilities. 1614. For one of your sharp-gap tasks, design a *measurement variant* that uses a more continuous metric: partial credit, edit distance, or a graded rubric. Re-evaluate. Does the apparent emergence survive the more continuous measurement? What does this tell you about whether the capability is "really" emergent? 162 163The point of this exercise is to internalise the emergence debate experientially. Reading about it is one thing; observing emergence (or its mirage) on tasks you designed is more direct. 164 165--- 166 167## The Current Frontier Models 168 169As of 2025, the frontier of language models is dominated by a small set of labs. The picture moves quickly, but the major players have been stable for several years. 170 171**OpenAI (GPT family).** GPT-4 (2023) was the model that established what "frontier" meant after GPT-3.5. GPT-4 was followed by GPT-4 Turbo (cheaper, faster) and GPT-4o (multimodal). The **o1** family (released 2024) introduced reasoning models (see Lesson 6.3) and represents a different paradigm than the standard chat models. The GPT-5 release (2025) consolidated the GPT and o-line into a unified family. OpenAI's models are closed-weight; available via API. 172 173**Anthropic (Claude family).** **Claude** 3 (2024) and Claude 3.5 Sonnet (2024) were the company's frontier offerings; Claude 3.7 Sonnet and Claude 4 followed. Anthropic's models are particularly noted for thoughtful refusal behaviour, Constitutional AI as a public de
173sign choice, and strong performance on reasoning and code tasks. Closed-weight; available via API. 174 175**Google DeepMind (Gemini family).** **Gemini** 1.5 Pro (2024), Gemini 2.0 (2024), Gemini 2.5 (2025). Google's models are particularly strong on multimodality (long-context image and video understanding) and have the largest context windows of the frontier models. Closed-weight; available via API. 176 177**Meta (Llama family).** Llama 3 (2024), Llama 3.1 (with 405B **parameter count** variants), Llama 3.2 (small variants), Llama 4 (2025). Meta's models are *open-weight*: the weights are released publicly, allowing anyone to run, fine-tune, or modify them. The Llama family has become the foundation of much of the open-source AI ecosystem. 178 179**Mistral.** Mistral Large, Mistral Small, the various Mixtral mixture-of-experts models. Mistral has staked out a position as a European frontier lab with a mix of open and closed offerings. The Mistral models have been competitive with the closed-API offerings on cost-adjusted basis for many tasks. 180 181**DeepSeek.** A Chinese lab whose **DeepSeek** R1 (2025) made a significant impact: a reasoning model that matched OpenAI's o-line on several benchmarks while being released open-weight. DeepSeek-V3 and the DeepSeek-R1 distilled variants extended the family. 182 183**Other notable participants.** Cohere (enterprise-focused, closed-weight), AI21 Labs (Jurassic family), various Chinese labs (Qwen from Alibaba, Yi from 01.AI, GLM from Zhipu), various smaller research efforts. The frontier itself is concentrated; the next tier behind the frontier is broad. 184 185The picture is moving fast enough that any specific list will be partially outdated within months. The structural picture is more stable: a handful of frontier labs with similar-but-differentiated offerings, an open-weight ecosystem dominated by Llama derivatives and increasingly by **DeepSeek** derivatives, and a long tail of more specialised models. 186 187--- 188 189## The Open-Weight Ecosystem in More Detail 190 191The open-weight side of the field deserves a closer look, because the ecosystem dynamics around open-weight models have substantial implications for capability diffusion, research practice, and policy. 192 193**The Llama foundation.** Meta's Llama family has been the foundation of much of the open-weight ecosystem since Llama 1 (2023). Llama 2 (2023) was the first open-weight release at frontier scale; Llama 3 (2024) and Llama 3.1 (with 405B **parameter** variants) extended the family and substantially closed the gap between open-weight and closed-API frontier models. Llama 3.2 brought small variants (1B, 3B) trained partly via distillation. Llama 4 (2025) consolidated the family further. The Meta team has continued to invest in open-weight releases despite ongoing internal and external debate about whether this is the right policy. 194 195**Mistral and the mixture-of-experts variant.** Mistral has staked out a position with both open and closed offerings. Their Mixtral models introduced mixture-of-experts architecture into the open-weight world: models that have a much larger total parameter count but activate only a fraction of the parameters per query, producing better performance per inference dollar. The Mixtral family demonstrated that architectural innovation could happen in the open-weight world rather than being concentrated at the frontier labs. 196 197**DeepSeek and the reasoning-model release.** **DeepSeek**'s R1 release (2025) was the open-weight reasoning model that demonstrated the **o1** paradigm could be replicated outside of OpenAI's closed pipeline. The R1 release, alongside DeepSeek-V3 and various distilled variants, was a substantial moment for the open-weight ecosystem: it showed that the latest paradigm shift was reproducible openly. The DeepSeek work also raised the legal-and-IP questions discussed in Lesson 5.6, because the training is reported to have included synthetic data generated by closed-API models. 198 199**The fine-tuning ecosystem.** A substantial wave of community-fine-tuned models has emerged on top of the open-weight foundations. Hugging Face's model hub now contains tens of thousands of fine-tuned variants of the major open-weight families. Most are forgettable; some have become genuinely useful for specific applications. The ecosystem has also produced specialised research tools (instruction-tuning datasets, preference datasets, evaluation harnesses) that benef
199it closed and open development alike. 200 201**Quantisation and accessibility.** The QLoRA-and-friends technical work covered in Lesson 5.6 has made it possible to run frontier-scale open-weight models on consumer hardware. Llama 3.1 8B can run on a phone; Llama 3.1 70B can run on a single high-end GPU. The accessibility has dramatically broadened who can experiment with these models, with implications both for legitimate research and for misuse. 202 203**The legal and policy backdrop.** The open-weight question is increasingly a regulatory matter. The EU AI Act has provisions that address the responsibilities of providers of general-purpose AI models, with some specific carve-outs for open-source releases. The US has had ongoing policy discussion about whether very large open-weight releases should be subject to additional requirements. The picture is unsettled and likely to change substantially over the next few years. 204 205**Capability diffusion and the safety concern.** The case against open-weight releases has crystallised around specific failure modes: open-weight models can be fine-tuned to remove safety guardrails (this has been demonstrated repeatedly, often quickly after a release); open-weight models can be deployed at scale without the lab's monitoring; specific dangerous-capability evaluations may be harder to enforce when the model is open. The case for open-weight releases has crystallised around: research access, broader capability diffusion, prevention of monopoly concentration, and the empirical observation that misuse so far has been less severe than the worst predictions. 206 207The honest summary: there is no consensus answer. The labs have made different choices for defensible reasons. The ecosystem has continued to grow on both sides. The decision-relevant question for any specific deployment has moved beyond "open vs closed" in the abstract. The question now is "which available model fits this use case, given its capabilities, its license, its operational requirements, and the risk profile of the deployment." 208 209For an annotator: the open-weight vs closed distinction matters because it shapes who is doing the work that produces the data. Closed-API models are trained at the frontier labs; their preference data is collected by lab-hired contractors. Open-weight models, after the initial release, are fine-tuned by countless other groups using their own preference data, their own constitutions, their own quality standards. The data pipelines diverge. The skills you develop on one kind of work are partially but not fully transferable to the other. 210 211--- 212 213## A Closer Look at Benchmark Contamination 214 215The contamination problem mentioned earlier deserves more detailed treatment, because it is one of the most consequential issues for reading current frontier-model claims. 216 217**What contamination means.** The model has seen the test items in its training data and is producing answers from memory rather than from learned capability. The reported **benchmark** score reflects memorisation rather than generalisation. 218 219**How contamination happens.** Public benchmarks are scraped into training corpora over time. The **MMLU** questions, for example, have been on the internet since 2020; any model trained with web-scraped data after that date has likely seen many of them. The contamination is rarely intentional (labs do not deliberately train on benchmarks), but the pretraining corpus is too large to systematically de-contaminate. 220 221**Specific evidence of contamination.** Several lines of evidence indicate that current frontier models have seen substantial fractions of public benchmarks. Models can sometimes reproduce specific test items verbatim if asked. Models perform substantially better on benchmark questions written before their training cutoff than on similar-difficulty questions written after the cutoff. The performance gap on contaminated-vs-uncontaminated subsets has been measured directly in several studies. 222 223**Mitigation strategies.** Various benchmarks now use mitigation: time-stamped test sets (questions written after a known training cutoff), hidden test items (the public benchmark uses a separate evaluation set), continuously refreshed evaluation sets (LiveCodeBench, for example). These are useful but do not fully solve the problem; new benchmarks become contaminated as soon as they are public. 224 225**The implication for reading scores.** A high score on a public **benchmark** with a 2022 release date and a 2024-trained model is much weaker evidence of capability than the same score on a benchmark released after the model's training cutoff. Practitioners who weight benchmark scores heavily should also weight the contamination question seriously. 226 227**The implication for evaluation work.** Designing contamination-resistant evaluations is a real research-and-engineering activity. The work involves writing fresh test items, keeping test items out of the public corpus, and validating that models have not seen the items via various technical checks. For an annotator who moves into evaluation work, contamination resistance is part of the methodology to know. 228 229The general principle: benchmark numbers are partial signals, contaminated benchmarks are weaker partial signals, and contamination-resistant evaluation matters for serious capability assessment. Treating any single benchmark score as the truth, without considering contamination, elicitation, and saturation, is one of the most common ways non-technical readers get misled by AI capability claims. 230 231--- 232 233## What Distinguishes Frontier Models 234 235The frontier models look superficially similar. They all chat, answer questions, and produce reasonable code; the differences between them are real but subtle. A short tour: 236 237**Reasoning depth.** Some models are stronger on multi-step logical reasoning than others. **Claude** 4 and the o-line are noted for stronger reasoning; **Gemini** and GPT-4 are competitive but have different specific strengths. Reasoning benchmarks (GPQA, MATH, ARC-AGI) discriminate the models more than general benchmarks like **MMLU**. 238 239**Code capability.** SWE-bench scores discriminate the models on agentic coding tasks. Claude Sonnet variants have led several SWE-bench rankings; the o-line and GPT-4 variants are competitive. Code-specialised models (DeepSeek-Coder, the Codestral family from Mistral) sometimes outperform general-purpose frontier models on coding tasks. 240 241**Multimodal capability.** Gemini's models have been particularly strong on long-form video and image understanding; GPT-4o brought voice and image input to OpenAI's chat surface; **Claude** 3 introduced image capabilities, which Claude 3.5 further improved. The differences are partly architectural and partly trained capability. 242 243**Context window size.** Gemini's 2 million-token context is the largest frontier-model **context window** as of 2025; the others are typically 128K-200K tokens. Whether the larger window is useful in practice depends on the application. 244 245**Refusal behaviour.** The models differ substantially in what they will and will not do. **Claude** is noted for thoughtful refusal that explains the reasoning; GPT-4 has more permissive defaults; the open-weight models can be re-trained to remove refusals entirely. 246 247**Personality and tone.** Each frontier model has a distinct voice that emerges from its training data, **RLHF** choices, and constitution. Users with sustained experience can identify which model produced a response from the voice alone. This is real product differentiation even when the capability surfaces are similar. 248 249**Latency and cost.** Production deployments care about how fast and how expensive each query is. The frontier models differ by order-of-magnitude on these dimensions, with smaller and quantised variants used for high-volume applications. 250 251The right way to read these differences: the frontier is a Pareto frontier across many dimensions, and different deployments have different fits with different points on the frontier. The "best model" question is not well-posed without specifying the application. For a given application, the best model is the one whose dimensional profile fits the application's needs. 252 253--- 254 255## Open-Weight vs Closed: The Debate 256 257One of the most consequential debates in the field is whether frontier model weights should be released openly. The two positions: 258 259**The closed-weight position.** Frontier models are powerful enough that releasing them publicly carries real risks: the weights can be fine-tuned to remove safety guardrails, used to generate at scale by malicious actors, or studied to find adversarial vulnerabilities that would not have been found otherwise. Closed-weight deployment via API gives the lab control over how the model is used. OpenAI and Anthropic have argued versions of this position; Google has staked out a more middle position. 260 261**The open-weight position.** Closed-weight models concentrate power in a handful of labs, prevent independent safety research, prevent academic research that requires access to model internals, and slow the diffusion of AI capability into broader society. Open-weight models enable a much larger research community, allow fine-tuning for specific use cases, and prevent any single lab from monopolising the capability. Meta, Mistral, and **DeepSeek** have argued versions of this position. 262 263Both positions have weight. The empirical record so far has not produced a clear answer. Open-weight Llama variants have been used both for legitimate research and (allegedly) for harmful purposes; the closed-weight frontier labs have done both important safety research and have made decisions that many in the broader community disagree with. The debate continues, with various jurisdictions now considering legislation that would clarify which considerations should be binding. 264 265For an annotator's understanding: knowing which models are open-weight and which are closed shapes which deployments and uses are likely. An open-weight model can be fine-tuned by anyone with the resources; a closed-weight model is constrained by the API provider's policies. Both have implications for the kind of work the model is used for and for what quality expectations are reasonable. 266 267--- 268 269## Inference-Time Scaling: The New Axis 270 271Until 2024, "scaling" mostly meant training-time scaling: more parameters, more data, more training compute. The **o1** family introduced a new axis: **test-time compute** scaling. The same trained model, given more compute per query (more tokens of internal reasoning, more sampling, more search), produces dramatically better results on hard problems. This is a different way to spend the compute budget, and it has changed how the field thinks about capability. 272 273The basic finding from the **o1** work: on a fixed-quality model, accuracy on hard math problems improves with the amount of internal **chain-of-thought** tokens generated before the final answer. The improvement is substantial, sometimes 30â50 percentage points on competition math problems, and it scales smoothly with the inference budget. More thinking time produces better answers. 274 275This has several implications for the current frontier picture. 276 277**Capability is a function of inference budget, not just model size.** A smaller model with more inference compute can sometimes match a larger model with less inference compute. This changes the cost calculus of deployment: for hard tasks, paying for more inference time per query may be cheaper than paying for a larger model. 278 279**The capability frontier is no longer just "the largest model."** With reasoning models, a deployment that uses extended inference time on a frontier model achieves capabilities that no smaller-or-faster setup can. The capability frontier is now defined by both model scale and inference scale. 280 281**Different deployments operate at different points on the inference-time curve.** Casual chat queries get fast, low-compute responses. Hard reasoning queries get slow, high-compute responses. The deployment routes queries to the appropriate inference budget. This is a more sophisticated deployment architecture than uniform per-query inference, and it is increasingly standard. 282 283**The empirical scaling laws are different.** The Kaplan-and-Chinchilla-style **scaling laws** are about training compute. The new scaling-laws work, much of it from OpenAI, Anthropic, and **DeepSeek**, is about inference compute, and the curves have different shapes. Some hard problems show extremely steep returns to inference compute; others show shallow returns. Knowing which class your problem is in helps with deployment planning. 284 285**The evaluation regime needs to change.** Benchmarks that report "accuracy at default elicitation" can substantially underestimate models if the inference budget is the binding constraint on capability. As discussed in Lesson 5.9, **capability elicitation** is a real eval-design issue, and the inference-time scaling story makes it more important. A reported benchmark score is now partly a statement about the inference budget used to produce it. 286 287For an annotator's purposes: knowing that capability has an inference-time dimension matters because it shapes what kinds of preferences the model can be trained to produce. A reasoning model can be trained on preference data that includes the model's own intermediate reasoning, in addition to the final response. This opens annotation work that was not possible before: labelling intermediate reasoning steps, or labelling whether a **chain-of-thought** trace is sound. The new annotation task has its own skill profile, distinct from final-response labelling. 288 289--- 290 291## A Worked Example: Reading a Frontier Model Release Announcement 292 293To bring the lesson's content together, walk through what a calibrated reading of a frontier model release announcement looks like in practice. 294 295**The announcement (composite, drawn from real recent releases).** A lab announces a new model. The announcement claims: 296 297- "Highest-ever score on **MMLU** at 92.4%" 298- "Strong reasoning capabilities: beats GPT-4 on MATH by 12 points" 299- "Multimodal: handles images, audio, and video" 300- "Long context: 1 million tokens" 301- "Available now at this price per token" 302 303A non-technical reader sees an impressive list of capabilities. A calibrated reader asks specific questions about each claim. 304 305**Claim 1: 92.4% on MMLU.** Is **MMLU** saturated? Yes; every frontier model scores 85%+ on standard MMLU. The claim of "highest-ever" is a 1â2 point improvement at the saturated end, which is not a meaningful capability difference. What MMLU subset was tested? Was the harder MMLU-Pro tested? Without those numbers, the headline is mostly marketing. *Calibrated reading: not strong evidence of substantial capability advance.* 306 307**Claim 2: 12 points over GPT-4 on MATH.** What elicitation? MATH scores depend heavily on whether reasoning is enabled, on **chain-of-thought** prompting, on best-of-N sampling, on inference compute budget. If the comparison is against GPT-4 with default prompting and the new model used reasoning-mode plus best-of-32, the comparison is not apples-to-apples. *Calibrated reading: depends on the elicitation methodology, which the announcement may or may not disclose. Worth investigating before believing.* 308 309**Claim 3: Handles images, audio, and video.** What does "handles" mean? Reading text in images is one capability; reasoning about complex visual scenes is another; understanding minute-long video segments is another still. The announcement is vague enough that the claim could be supporting any of these or just the simplest. *Calibrated reading: not specific enough to evaluate. Look for specific multimodal benchmark results before drawing conclusions.* 310 311**Claim 4: 1 million token context window.** What is the model's "lost-in-the-middle" performance at 500K-token context length? Long **context window**s are sometimes nominal: the model can technically accept the input but performance degrades severely beyond a certain length. Independent evaluations of long-context performance often differ substantially from the nominal context window. *Calibrated reading: capability claim depends on long-context benchmark performance, not the nominal window size.* 312 313**Claim 5: Pricing.** Price per token is real and measurable. Worth comparing agains
313t competitors at the same capability tier, accounting for the **test-time compute** story above. A more capable model with higher per-token cost may be the right choice for hard tasks; a cheaper less-capable model may be the right choice for routine ones. *Calibrated reading: the most concrete claim, easiest to verify.* 314 315**The synthesis.** Of the five claims, only the price is fully verifiable from public information. The capability claims are partly indicative but depend heavily on how they were elicited and measured. The right response is calibrated interest: *this is potentially a real advance, here are the specific things I want to see before being more confident*, rather than either dismissal or full acceptance. 316 317This is the calibrated-reading skill in action; it is care rather than cynicism. The labs producing announcements have legitimate reasons to highlight their best results, and those results are often genuinely good. But the announcement format compresses real information; reading carefully and knowing what additional information would resolve specific questions is what separates serious engagement from hype-following. 318 319For a Sovrano annotator interested in how the field's work gets communicated publicly: announcements like these are the public-facing tip of substantial internal evaluation work. The numbers come from labs with internal evaluation teams running dozens of benchmarks; the announcement selects the most flattering subset. Senior annotators who eventually move into evaluation roles encounter both sides (the comprehensive internal evaluation and the selective external announcement), and the contrast is informative about how the field communicates with itself versus with the broader public. 320 321--- 322 323## Reading Benchmark Results 324 325A practical literacy worth developing: how to read **benchmark** results without being misled by them. 326 327**The benchmark might be saturated.** If every frontier model scores 90%+, the benchmark is saturated and the differences between models are not meaningful at the headline level. Look for sub-scores or harder versions (MMLU-Pro vs **MMLU**; GPQA-Diamond vs GPQA). 328 329**The benchmark might be contaminated.** Public benchmarks are scraped into training data over time. A model trained with data through mid-2024 has likely seen a substantial fraction of public 2022-era benchmark questions. The headline number on these benchmarks should be discounted. 330 331**The elicitation matters.** As discussed in Lesson 5.9, benchmark scores depend on how the model was prompted. A score reported with "best-of-N sampling at N=5 with a scaffolded prompt" is not the same as a score reported with "default prompt, single attempt." Compare scores at the same elicitation level. 332 333**The benchmark might not match deployment.** A high **MMLU** score does not predict whether the model is useful for the task you actually care about. Deployment-relevant evaluation is usually domain-specific and has to be done separately from headline benchmarks. 334 335**The benchmark might be optimised for.** Labs sometimes optimise specifically for scoring well on standard benchmarks. The optimisation can be subtle (training on similar problems) or overt (specifically targeting the benchmark in evaluation runs). High scores on a benchmark are a partial signal of capability; perfect scores often indicate optimisation rather than genuine capability gain. 336 337The HELM framework (Liang et al., Stanford CRFM, 2022) was an early attempt to address some of these issues by evaluating models across many dimensions simultaneously. It has not fully replaced single-benchmark comparison, but it is a useful complement when reading frontier-model claims. 338 339The general advice: treat any single benchmark score as a partial signal, look for cross-benchmark consistency, prefer evaluations that match your specific use case, and weight recent independent evaluations more than lab self-reports. 340 341--- 342 343## Practice Exercise 6.2.2 344*Allow 15â20 minutes.* 345 346You are reviewing a frontier-lab announcement claiming a new state-of-the-art result. The announcement reports several benchmark numbers and claims the model outperforms competitors on most tasks. Spend at least twenty minutes: 347 3481. Pick a real recent frontier-model release announcement. Identify all the specific benchmark numbers it reports. 3492. For each benchmark, determine: what does it measure? Is it saturated? Is contamination a concern? What elicitation methodology was used? Is the methodology comparable to what competitors used? 3503. Pick one benchmark where the announcement claims a substantial improvement. Find at least one independent third-party evaluation of the same model on the same benchmark. Compare. Are the numbers consistent? If not, why might they differ? 3514. Write a one-paragraph "calibrated reading" of the announcement: what claims should be taken as solid evidence, what claims should be discounted, and what claims you would want more information about before believing. 352 353The point of this exercise is to practise the calibrated-reading skill. Frontier-lab announcements have implications beyond the labs themselves; they shape policy debates, market dynamics, and public perception of AI capability. Reading them with appropriate scepticism is a literacy that compounds over time. 354 355--- 356 357## What "State of the Art" Actually Means 358 359The phrase "state of the art" appears in nearly every frontier-model announcement. It is worth unpacking what the phrase actually means in 2025, because the meaning has drifted from what it meant in earlier eras. 360 361In 2018â2020, "state of the art" on a **benchmark** usually meant a single specific number: the highest score any reported method had achieved on the benchmark. The number was comparable across methods because everyone evaluated on the same test set with broadly similar elicitation strategies. 362 363That era is over. In 2025, "state of the art" has at least four distinct meanings, often conflated. 364 365**Best on a single specific benchmark.** A model achieves the highest score on, say, GPQA-Diamond. This is the older meaning, and it is still meaningful when the benchmark is well-designed and not saturated. But it is one number, and the elicitation methodology matters substantially. Reading "state of the art on GPQA" requires asking whether the elicitation was comparable to prior reports. 366 367**Best on a basket of benchmarks.** A model is at or near the top on many benchmarks simultaneously. This is a more robust claim than single-benchmark dominance, but it has its own issues: a model can be optimised to perform well on a basket of standard benchmarks without being deployment-grade on any specific application. 368 369**Best at a deployment-relevant capability.** A model is the most reliable for a specific application: coding, reasoning, multilingual translation, long-form writing. This is the meaning most useful for practitioners, but it is not what "state of the art" usually means in announcements. The deployment-relevant capability is often only weakly correlated with the headline benchmark numbers. 370 371**Most recent release.** Sometimes "state of the art" just means "the newest model from a frontier lab." This is a tautological claim (the newest is the newest), but it gets used in marketing because the newest is usually at least as good as the previous generation. 372 373The drift is partly because the field has more models than benchmarks can usefully discriminate. Multiple frontier models score within a few percentage points on most standard benchmarks; the differences are within measurement noise on many comparisons; the headline "state of the art" claim is partly a function of which benchmarks the announcement chose to highlight rather than which model is genuinely best. 374 375A useful framing: rather than "state of the art," ask "best at *what*?" The answer specifies the application, the benchmark, and the elicitation. Without those specifications, the claim is mostly marketing rather than
375information. 376 377For an annotator's purposes: the drift in what "state of the art" means matters because the data the team is producing might be used to push a specific dimension of the capability frontier rather than the whole capability surface. A team that is collecting preference data on math problems is targeting math-frontier capability; a team collecting preference data on creative writing is targeting writing-frontier capability. Knowing which dimension your work targets is part of understanding what your contribution does. 378 379--- 380 381## A Closer Look at Specific Frontier Models 382 383The model-by-model summary above gave the headline picture. A closer look at what distinguishes the major models in 2025 is worth providing for context. 384 385**OpenAI's GPT-5 family.** OpenAI's flagship combines the GPT-line's strong general capability with the **o1**-line's reasoning paradigm. Specific strengths: strong on multimodal queries, strong on reasoning when invoked in reasoning mode, strong on coding via the dedicated code variants. Specific weaknesses: closed-API only, comparatively expensive at the frontier tier, less transparent about training methodology than some competitors. The platform deployments (ChatGPT, the API) reach hundreds of millions of users. 386 387**Anthropic's Claude 4 family.** **Claude** is noted for its conversational quality, thoughtful refusal behaviour, and Constitutional AI grounding. Claude 4 Opus and Sonnet variants are competitive with OpenAI's frontier on most benchmarks; the Sonnet variants have been particularly strong on coding (leading SWE-bench rankings at various points). Specific strengths: clear and consistent persona, good handling of nuanced ethical edge cases, strong long-context handling. Specific weaknesses: tighter refusal patterns sometimes block legitimate requests, smaller user-facing platform than OpenAI. 388 389**Google's Gemini 2.5 family.** **Gemini**'s multimodal foundation is the strongest in the frontier set; Gemini's context windows are the largest. Gemini 2.5 has reasoning capabilities competitive with the o-line. Specific strengths: long-context handling, multimodal depth, integration with Google's broader product surface. Specific weaknesses: somewhat inconsistent on conversational quality, refusal behaviour has shifted across versions in ways that have frustrated users. 390 391**Meta's Llama 4 family.** Meta's open-weight flagship. Llama 4's largest variants are competitive with closed-API frontier models on most benchmarks; the smaller variants are the best open-weight options for their size. The open-weight release is the distinguishing feature; the model can be fine-tuned, deployed locally, and modified in ways that closed-API models cannot. Specific strengths: open weights, ecosystem support, decreasing gap to closed-API frontier. Specific weaknesses: lab-supplied tooling is less polished than the closed-API providers, distillation-derived smaller variants sometimes inherit personality features from the larger Llama models. 392 393**DeepSeek's R1 and V3 families.** Open-weight reasoning model and general-purpose model. **DeepSeek**'s release pattern of high-quality open-weight models has substantially shifted the landscape; the R1 release in particular demonstrated that the reasoning-model paradigm could happen openly. Specific strengths: open-weight reasoning capability, aggressive pricing for hosted versions, contributions to the open ecosystem. Specific weaknesses: the legal-and-IP question discussed in Lesson 5.6 is unresolved for some of the training methodology, and policy reactions have varied across jurisdictions. 394 395**Mistral's Large and Small families.** European frontier lab with mixed open and closed releases. Mistral's mixture-of-experts variants (Mixtral) introduced the architecture into the open ecosystem; Mistral Large is a closed-API frontier offering. Specific strengths: European regulatory positioning, MoE architecture, strong multilingual capability. Specific weaknesses: smaller scale of operations than the US labs, less prominent positioning in English-dominated discourse. 396 397The picture across these: each frontier lab has a distinctive position, and the differences are real but smaller than the headline announcements suggest. Knowing which model is the right fit for a specific task is increasingly an empirical question: different models are better at different specific things, and the answer depends on your task. 398 399For a deployment decision, the practical advice: evaluate multiple frontier models on your specific use case, weight independent evaluations more than lab self-reports, and prefer models whose strengths match your task even if they are not the headline-leader on general benchmarks. The "best frontier model" question is not well-posed; the "best frontier model for this specific application" question is. 400 401--- 402 403## What Reading the Frontier Requires in Practice 404 405Reading the frontier landscape with the discipline this lesson has been pushing for is not a one-time skill. It is something a practitioner does continuously, because the frontier moves continuously and the calibration of any specific claim has to be redone every few months. The discipline is partly intellectual (asking the right questions about new claims) and partly social (knowing which voices in the field are reliable analysts versus which are marketing-adjacent commentators). 406 407--- 408 409## A Closing Note on Reading the Frontier 410 411This lesson has covered both **emergence** as a phenomenon and the current frontier-model landscape. The two threads connect: emergence is what produced the current frontier, and the current frontier is where the next round of emergence (or apparent emergence, or genuine new capabilities) will be observed. 412 413The practical literacy the lesson aims to build is *calibrated reading of frontier claims*. When a lab announces a new capability, the calibrated reader asks: is this genuinely new, or is it a marketing rephrasing of an existing capability? Is the **benchmark** saturated? Is the elicitation comparable? What does an independent evaluation say? Is the claim domain-specific or general? Where would the limits of the claimed capability show up? 414 415This kind of reading is increasingly important because frontier-lab announcements have implications beyond the labs themselves: they shape policy debates, market dynamics, and public perception of AI. Reading them as marketing-with-substance, rather than as either marketing-only or substance-only, is the disciplined posture. 416 417For an annotator: the frontier landscape this lesson covers is the context in which your work happens. The labs you might work with, the deployments your work might inform, and the evaluations your output might contribute to all sit within the structure this lesson sketched. Knowing the structure helps you place your contribution and helps you reason about where the most valuable work is happening. 418 419--- 420
421## Connections 422 423This lesson sits between Lesson 6.1 (which covered the deployed-model capability surface in general) and Lesson 6.3 (which covers the latest extensions: reasoning models, multimodality, **RAG**, tool use). Together, the three lessons form a survey of where current AI capability is and how to think about it. 424 425Backward, the lesson connects to Lesson 2.5 (early scaling behaviour, where the patterns that produced emergence were first observed) and to Series 5 (the training pipeline that produced the current frontier models). Without Series 5, the differences between frontier models look mysterious; with it, they look like predictable consequences of specific training-pipeline choices. 426 427Forward, the lesson connects to Series 7 (Red Teaming, which probes the failure modes of these specific frontier models) and Series 8 (which extends the analysis to capabilities that are still emerging: agentic systems, world models, reasoning at scale). 428 429--- 430 431## What Annotators Should Know About the Frontier Landscape 432 433The frontier-model landscape this lesson has surveyed shapes annotation work in several specific ways. 434 435**Different labs run different annotation pipelines.** The labs covered above each have their own annotation operations, with their own guidelines, their own quality standards, their own contractor pools. Anthropic's helpful-and-harmless framing produces different preference data than OpenAI's preparedness framework. Google's product-integration focus produces different demonstrations than Meta's open-source focus. An annotator who w
435orks across labs (or who follows the lab's preferences) is implicitly working across distinct value frameworks. 436 437**The annotation cycle is now part of competitive dynamics.** Frontier labs compete on how fast they can identify and address failure modes in their deployed models. A lab with a faster annotation pipeline can ship improvements more quickly. The pressure to scale and accelerate annotation is real, and it has implications for annotation quality (faster is sometimes worse) and for annotator working conditions. 438 439**Open-weight models bring annotation work to a much broader community.** Open-weight model fine-tuning happens not just at frontier labs but at startups, research groups, and individual practitioners. The annotation skills covered in this curriculum are valuable in all of these contexts, with varying compensation profiles. The career path now extends beyond the frontier labs to the broader fine-tuning ecosystem. 440 441**Reasoning models add new annotation tasks.** As discussed above, the **o1**-style paradigm means there is now annotation work on reasoning-step quality, on **chain-of-thought** validity, on process-level evaluation rather than just final-response evaluation. This is a new specialisation, and the skills overlap with but are not identical to standard preference annotation. 442 443**Multimodal annotation is a distinct skillset.** Annotating image, audio, and video preferences requires different cognitive work than annotating text. Image-quality annotators need a different eye than text-quality annotators; audio annotators need to listen carefully; video annotators have to follow temporal dynamics. The multimodal frontier brings new annotation specialisations. 444 445**Domain expertise is increasingly valuable.** As deployments specialise into specific domains (medicine, law, finance, code), the value of domain-expert annotators grows. A general annotator can produce general-purpose preferences; a medical-expert annotator can produce medical-specific preferences that the general annotator cannot. The frontier-deployment economy increasingly needs both. 446 447**Evaluation work is becoming a profession.** The frontier labs and the independent evaluators (METR, Apollo, AISI) are creating a profession of "AI evaluator": people whose primary work is rigorously assessing model outputs. The skill set overlaps substantially with senior annotation work, and the career pipeline from senior annotator to professional AI evaluator is real. 448 449For an annotator thinking about where to invest skill development: domain expertise, evaluation methodology, multimodal capability, and reasoning-step assessment are all areas where demand is growing faster than supply. Each is a specialisation that can grow alongside the frontier rather than being commoditised by automation. 450 451--- 452 453## Key Takeaways 454 455- **Emergent capabilities** appear at scale without being explicitly trained. Wei et al. (2022) catalogued examples; Schaeffer et al. (2023) argued that some apparent emergence is a measurement artefact. The current consensus: some capabilities are genuinely emergent, others look more emergent than they are. 456- **In-context learning** (the ability to learn from examples in the prompt without fine-tuning) is the most commercially important emergent capability. It is what makes prompting a coherent discipline. 457- The current frontier consists of a small set of labs: OpenAI, Anthropic, Google DeepMind, Meta, Mistral, **DeepSeek**, and a few others. The frontier moves fast but the structural picture is stable. 458- Frontier models differ on **reasoning depth, code capability, multimodality, context window, refusal behaviour, personality, and latency/cost**. The "best model" is application-dependent. 459- The **open-weight vs closed-weight debate** is unresolved and increasingly consequential for AI policy. Both positions have weight; the empirical record so far has not produced a clear answer. 460- Reading **benchmark** results requires care: saturation, contamination, elicitation differences, and benchmark optimisation all distort headline numbers. Triangulate across benchmarks, prefer independent evaluations, and match benchmarks to deployment use cases. 461 462--- 463 464## A Note on the Pace of Change 465 466A practical caveat: the model landscape is moving fast enough that any specific list of frontier models and capabilities is partially out of date by the time it is read. The version of this lesson written in 2025 will be missing several frontier models that exist by 2026, and will describe some current models in ways that no longer reflect the latest releases. 467 468The structural picture should be more durable: a small number of frontier labs, a growing open-weight ecosystem, multiple axes of capability scaling, and an expanding set of specialisations. The specific models named here are illustrative rather than load-bearing. 469 470For a current practitioner, keeping up with the landscape requires active tracking. Some useful sources: the lab-internal technical reports (released for major model launches), the LMSys ChatBotArena rankings (live user-facing comparison), the various benchmark leaderboards (LMSYS, OpenLLM Leaderboard for open models), and the newsletter ecosystem covering AI developments (Interconnects from Nathan Lambert, Import AI from Jack Clark, the Stratechery analysis from Ben Thompson). No single source is sufficient; tracking the landscape requires triangulating across several. 471 472The pace itself is a feature of the field as much as any specific model is. Working in AI in 2025 means working in a context where the relevant tools and capabilities are different from what they were six months earlier, and will be different again six months later. Comfort with this rate of change is part of the discipline. Annotators who stay engaged with the landscape (reading announcements, tr
472ying new models, noticing new capabilities) are positioned to grow alongside the field; annotators who treat their initial training as the complete picture are not. 473 474--- 475 476## Open Questions in the Current Landscape 477 478A few open questions that the current frontier landscape leaves unresolved are worth flagging for completeness. 479 480**Will the open-weight gap continue to close, stabilise, or widen?** The Llama and **DeepSeek** lines have closed substantial ground on the closed-API frontier in 2024â25. Whether this continues depends on factors that are partly technical (how much training the open-weight ecosystem can sustain) and partly policy (whether regulation encourages or discourages open releases). Both outcomes are possible. 481 482**Will reasoning-model capabilities generalise or stay specialised?** The **o1** family is dramatically better at math and code; it is not as much better on open-ended creative work or social-reasoning tasks. Whether the reasoning-model paradigm will extend to these other domains, or whether it is structurally limited to verifiable-answer tasks, is an open empirical question. 483 484**Will the frontier converge or diverge across labs?** Frontier models from different labs are currently quite similar in headline benchmarks but distinguishable in personality and specific capability strengths. Whether this differentiation will deepen (with each lab carving out a niche) or fade (with frontier capability becoming a commodity) is uncertain. 485 486**How will multimodal capability evolve?** The current **vision-language model** paradigm works but has clear limits. Whether the next generation of multimodal models will look architecturally similar (just bigger, better trained) or architecturally different (new approaches to grounding language in non-textual modalities) is open. 487 488**What will the regulatory environment look like?** The EU AI Act is partially operational; similar legislation is at various stages elsewhere. Whether the regulatory environment will favour closed-API deployment, open-weight ecosystems, or some specific compromise is consequential and unresolved. 489 490**Will the cost trajectory continue?** Per-token cost for frontier-quality output has fallen by orders of magnitude over the last several years. Whether this continues (driven by competition, hardware improvements, and inference optimisation) or stabilises is itself uncertain. The cost trajectory shapes which deployments are economically viable. 491 492These open questions are not blocking issues for current work, but they shape the trajectory of the field. A practitioner who is aware of the open questions can make decisions that are robust to multiple outcomes, rather than betting heavily on a specific outcome that may not occur. 493 494--- 495 496## Further Reading 497 498- **Wei, Jason et al., *Emergent Abilities of Large Language Models* (Google Research, Transactions on Machine Learning Research, 2022).** The foundational emergence paper. Read further than the introduction; the catalogue of specific emergent capabilities is the substantive content. 499- **Schaeffer, Rylan et al., *Are Emergent Abilities of Large Language Models a Mirage?* (arXiv, 2023).** The counter-argument. Read in conjunction with Wei et al. for the full debate. 500- **Liang, Percy et al., *Holistic Evaluation of Language Models* (Stanford CRFM, arXiv, 2022).** The HELM framework. A more sophisticated approach to benchmarking than single-metric comparison. 501- **Brown, Tom B. et al., *Language Models are Few-Shot Learners* (OpenAI, 2020).** The original GPT-3 paper, which demonstrated in-context learning at scale. Worth reading for the historical record. 502- **Dubey, Abhimanyu et al., *The Llama 3 Herd of Models* (Meta, 2024).** The detailed technical report on Llama 3. The most comprehensive public account of an open-weight frontier-model training pipeline. 503- **Anthropic, OpenAI, Google DeepMind, and Meta technical reports on their frontier-model releases (various years).** Worth reading the introduction sections for each, since they reveal the framing each lab uses for its own work. 504`;export{n as default};
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.