AI Infra
0%
References

References

AuthorChangkun Ou
Reading time~186 min
[119th United States Congress 2025] 119th United States Congress. GENIUS act (S.1582): Guiding and establishing national innovation for U.S. stablecoins act. https://www.congress.gov/bill/119th-congress/senate-bill/1582
[Abadi et al. 2016] Abadi, Chu, Goodfellow, McMahan, Mironov, Talwar, Zhang. Deep learning with differential privacy. https://arxiv.org/abs/1607.00133
Abadi et al. introduce DP-SGD, training deep neural networks under differential privacy via per-example gradient clipping, Gaussian noise, and a moments accountant for tighter privacy budget tracking.
[Abadi et al. 2016] Abadi, Barham, Chen, Chen, Davis, Dean, Devin, Ghemawat, Irving, Isard, Kudlur, Levenberg, Monga, Moore, Murray, Steiner, Tucker, Vasudevan, Warden, Wicke, Yu, Zheng. TensorFlow: A system for large-scale machine learning. https://arxiv.org/abs/1605.08695
Google's second-generation system: a static dataflow graph executed by a session, partitionable across hundreds of heterogeneous machines, rebuilt from the lessons of DistBelief.
[Abbas et al. 2023] Abbas, Tirumala, Simig, Ganguli, Morcos. SemDeDup: Data-efficient learning at web-scale through semantic deduplication. https://arxiv.org/abs/2303.09540
SemDeDup uses embeddings from pre-trained models to identify and remove semantically similar but non-identical duplicates, cutting web-scale training data by 50
[Abhyankar et al. 2025] Abhyankar, Qi, Zhang. OSWorld-human: Benchmarking the efficiency of computer-use agents. arXiv preprint arXiv:2506.16042. https://arxiv.org/abs/2506.16042
Measures what success rates hide: tens of minutes per task, 2.7 to 4.3 times more steps than the human-optimal trajectory, and latency dominated by the model calls rather than the actions.
[Acun et al. 2021] Acun, Murphy, Wang, Nie, Wu, Hazelwood. Understanding training efficiency of deep learning recommendation models at scale. https://arxiv.org/abs/2011.05497
[Agrawal et al. 2024] Agrawal, Kedia, Panwar, Mohan, Kwatra, Gulavani, Tumanov, Ramjee. Taming throughput-latency tradeoff in LLM inference with sarathi-serve. https://arxiv.org/abs/2403.02310
Sarathi-Serve introduces chunked-prefill and stall-free batching to eliminate the throughput-latency tradeoff in LLM inference, achieving up to 5.6x higher serving capacity over vLLM.
[Agrawal 2026] Agrawal. The economics of generative AI: Two years later. https://apoorv03.com/p/the-economics-of-generative-ai-two
Agrawal estimates that in 2026 semiconductors still capture about 79 percent of AI gross profit against 13 percent for infrastructure and 8 percent for applications, down from roughly 87/9/4 two years earlier, and argues value will migrate up the stack as in past compute supercycles, but slowly.
[Ahmadian et al. 2024] Ahmadian, Cremer, Gallé, Fadaee, Kreutzer, Pietquin, Üstün, Hooker. Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in llms. arXiv preprint arXiv:2402.14740. https://arxiv.org/abs/2402.14740
Simple REINFORCE-style policy gradient (RLOO) outperforms PPO and DPO for RLHF alignment of LLMs, with lower compute cost and no need for PPO's actor-critic complexity.
[AI News 2026] AI News. Did meta sacrifice its open-source identity for a competitive AI model?. https://www.artificialintelligence-news.com/news/meta-muse-spark-ai-model-open-source/
[Ainslie et al. 2023] Ainslie, Lee-Thorp, Jong, Zemlyanskiy, Lebron, Sanghai. GQA: Training generalized multi-query transformer models from multi-head checkpoints. https://aclanthology.org/2023.emnlp-main.298/
GQA introduces grouped-query attention, an interpolation between MHA and MQA that matches MHA quality while approaching MQA inference speed, and provides a recipe to uptrain existing MHA checkpoints using only 5
[Alayrac et al. 2022] Alayrac, Donahue, Luc, Miech, Barr, Hasson, Lenc, Mensch, Millican, Reynolds, others. Flamingo: a visual language model for few-shot learning. https://arxiv.org/abs/2204.14198
Flamingo is a Visual Language Model (VLM) family that bridges frozen vision and language models with a Perceiver Resampler and gated cross-attention, enabling few-shot learning across 16 image and video understanding tasks.
[Albergo and Vanden-Eijnden 2023] Albergo, Vanden-Eijnden. Building normalizing flows with stochastic interpolants. https://arxiv.org/abs/2209.15571
This paper introduces stochastic interpolants, a framework for building continuous normalizing flows between any two densities via a simple quadratic loss that avoids backpropagation through ODE solvers.
[Albergo et al. 2023] Albergo, Boffi, Vanden-Eijnden. Stochastic interpolants: a unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797. https://arxiv.org/abs/2303.08797
This paper introduces stochastic interpolants, a unified framework for generative modeling that subsumes flow-based and diffusion-based methods by bridging any two densities exactly on a finite time interval via ODE or SDE dynamics.
[Alphabet 2025] Alphabet. Sundar pichai on alphabet Q3 2025 earnings (token volume). https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q3-2025/
Alphabet's disclosed monthly token-processing volume grew roughly twentyfold year over year through 2025, the demand-side evidence against the pure vendor-financing critique.
[Alsup 2025] Alsup. Bartz v. anthropic PBC, order on fair use, no. 3:24-cv-05417 (N.D. cal.). https://docs.justia.com/cases/federal/district-courts/california/candce/3:2024cv05417/434709/231
Judge Alsup holds that training LLMs on lawfully acquired books is exceedingly transformative fair use, while downloading and permanently retaining millions of pirated shadow-library books is not.
[Amatriain and Basilico 2012] Amatriain, Basilico. Netflix recommendations: Beyond the 5 stars (part 1). https://netflixtechblog.com/netflix-recommendations-beyond-the-5-stars-part-1-55838468f429
The insiders' account of the Netflix Prize's production aftermath: only two component algorithms of the winning ensemble were ever deployed, because the offline accuracy gain did not justify the engineering cost and the product had moved from rating prediction to ranking.
[Amazon 2025] Amazon. Amazon Form 10-Q (server useful-life reduction). https://www.sec.gov/Archives/edgar/data/1018724/000101872425000036/amzn-20250331.htm
Amazon shortened a subset of servers and networking equipment from six years back to five, effective 2025, citing the increased pace of technology development in AI and machine learning.
[Amazon Web Services 2025] Amazon Web Services. Firecracker snapshotting. https://github.com/firecracker-microvm/firecracker/blob/main/docs/snapshotting/snapshot-support.md
Firecracker's snapshot-support docs describe how to serialize a running microVM to disk and restore it with on-demand memory loading via MAP_PRIVATE, enabling fast workload cloning and resume.
[AMD 2020] AMD. AMD SEV-SNP: Strengthening VM isolation with integrity protection and more. https://www.amd.com/system/files/TechDocs/SEV-SNP-strengthening-vm-isolation-with-integrity-protection-and-more.pdf
The whitepaper for VM-level confidential computing: encrypt and integrity-protect a whole virtual machine against a malicious hypervisor, so unmodified stacks can run confidentially.
[AMD 2025] AMD. AMD and OpenAI announce strategic partnership (form 8-K exhibit). https://www.sec.gov/Archives/edgar/data/2488/000119312525230895/d28189dex991.htm
A commitment for 6GW of AMD Instinct GPUs alongside a warrant granting OpenAI up to  10
[Ameisen et al. 2025] Ameisen, Lindsey, Pearce, Gurnee, Turner, Chen, Citro, Abrahams, Carter, Hosmer, Marcus, Sklar, Templeton, Bricken, McDougall, Cunningham, Henighan, Jermyn, Jones, Persic, Qi, Thompson, Zimmerman, Rivoire, Conerly, Olah, Batson. Circuit tracing: Revealing computational graphs in language models. https://transformer-circuits.pub/2025/attribution-graphs/methods.html
Circuit Tracing introduces attribution graphs that reveal step-by-step computational mechanisms in language models by replacing MLPs with interpretable cross-layer transcoders.
[Amershi et al. 2019] Amershi, Weld, Vorvoreanu, Fourney, Nushi, Collisson, Suh, Iqbal, Bennett, Inkpen, Teevan, Kikin-Gil, Horvitz. Guidelines for human-AI interaction. https://www.microsoft.com/en-us/research/publication/guidelines-for-human-ai-interaction/
This CHI paper distills eighteen generally applicable guidelines for user-facing AI products and validates them with practitioners reviewing AI-infused products.
[Anderson 1982] Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications 12(3):313–326. https://doi.org/10.1016/0304-4149(82)90051-5
Anderson (1982) derives the reverse-time SDE for a forward diffusion process, showing the reverse drift depends on the score function of the marginal distribution.
[Anderson 2025] Anderson. Don't build an RL environment startup. https://benanderson.work/blog/dont-build-rl-env-startup/
A practitioner blog post arguing that selling RL training environments to frontier AI labs is a structurally poor business because model capabilities quickly devalue each environment to zero.
[Android Developers 2026] Android Developers. Announcing gemma 4 in the aicore developer preview. https://android-developers.googleblog.com/2026/04/AI-Core-Developer-Preview.html
[Anil et al. 2024] Anil, Durmus, Sharma, others. Many-shot jailbreaking. https://www.anthropic.com/research/many-shot-jailbreaking
Many-shot jailbreaking exploits large context windows by packing hundreds of faux harmful-dialogue demonstrations into a single prompt, overriding LLM safety training via in-context learning.
[Ankner et al. 2024] Ankner, Parthasarathy, Nrusimha, Rinard, Ragan-Kelley, Brandon. Hydra: Sequentially-dependent draft heads for medusa decoding. https://arxiv.org/abs/2402.05109
Hydra heads are sequentially-dependent draft heads for speculative decoding that condition each prediction on prior candidate tokens, improving throughput by up to 2.70x over autoregressive decoding.
[Ansel et al. 2024] Ansel, Yang, He, Gimelshein, Jain, Voznesensky, Bao, Bell, Berard, Burovski, Chauhan, Chourdia, Constable, Desmaison, DeVito, Ellison, Feng, Gong, Gschwind, Hirsh, Huang, Kalambarkar, Kirsch, Lazos, Lezcano, Liang, Liang, Lu, Luk, Maher, Pan, Puhrsch, Reso, Saroufim, Siraichi, Suk, Zhang, Suo, Tillet, Zhou, Wang, Zou, Wang, Mathews, Wen, Chanan, Wu, Chintala. PyTorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation. https://docs.pytorch.org/assets/pytorch2-2.pdf
The retrospective on TorchScript's failure and the design of torch.compile: capture graphs from Python bytecode, break the graph where capture fails, and preserve eager semantics by construction.
[Anthropic 2023] Anthropic. Claude's constitution. https://www.anthropic.com/news/claudes-constitution
Anthropic describes Claude's constitution as an explicit and editable set of principles used to guide harmlessness and helpfulness through Constitutional AI.
[Anthropic 2024] Anthropic. Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku. https://www.anthropic.com/news/3-5-models-and-computer-use
The first frontier model to offer computer use in public beta: general computer skills over per-task tools, released with the plain warning that it was experimental, cumbersome, and error-prone.
[Anthropic 2024] Anthropic. Introducing contextual retrieval. https://www.anthropic.com/news/contextual-retrieval
Contextual retrieval prepends an LLM-generated situating sentence to each chunk before embedding and BM25 indexing, cutting the top-20 retrieval failure rate by 49
[Anthropic 2024] Anthropic. Model context protocol. https://modelcontextprotocol.io
MCP is an open standard for connecting AI applications and agents to external systems such as data sources, tools, and workflows through a common integration layer.
[Anthropic 2025] Anthropic. Prompt caching. https://platform.claude.com/docs/en/build-with-claude/prompt-caching
Anthropic's prompt-caching documentation prices cached input reads at one tenth of the base input token rate, with cache writes billed at a premium, making prompt structure an explicit cost lever for API callers.
[Anthropic 2025] Anthropic. Claude code: Checkpointing and session forks. https://code.claude.com/docs/en/checkpointing
Claude Code's checkpointing reference page documents how to track, rewind, and summarize Claude's file edits and conversation to manage session state.
[Anthropic 2025] Anthropic. Memory tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool
Anthropic's memory tool lets Claude store and retrieve information across conversations as a client-side directory of memory files it views, creates, and edits, keeping active context focused while knowledge persists between sessions.
[Anthropic 2025] Anthropic. Bringing memory to claude. https://claude.com/blog/memory
Memory scoped per project with a user-visible, editable summary and incognito chats, released with the stated safety testing of whether memory could reinforce harmful conversational patterns.
[Anthropic 2025] Anthropic. Piloting claude for chrome. https://www.anthropic.com/news/claude-for-chrome
A browser agent piloted with published red-team results: safeguards cut prompt-injection attack success from roughly a quarter of attempts to a tenth, framed by the vendor as not yet enough for wide deployment.
[Anthropic 2025] Anthropic. How we built our multi-agent research system. https://www.anthropic.com/engineering/multi-agent-research-system
Anthropic's research system uses an orchestrator-worker pattern: a lead Claude Opus 4 agent coordinating parallel Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2
[Anthropic 2025] Anthropic. Code execution with MCP: Building more efficient AI agents. https://www.anthropic.com/engineering/code-execution-with-mcp
Anthropic's engineering post shows that pairing MCP with an in-process code execution environment reduces agent token consumption, latency, and round-trips by handling filtering, control flow, and state persistence in code rather than via repeated tool calls.
[Anthropic 2025] Anthropic. Equipping agents for the real world with agent skills. https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
Agent Skills package instructions, scripts, and resources as a directory with a SKILL.md whose name and description are the only tokens held in context; the full contents load on demand, progressive disclosure as a first-class mechanism.
[Anthropic 2025] Anthropic. Open-sourcing circuit tracing tools. https://www.anthropic.com/research/open-source-circuit-tracing
[Anthropic 2025] Anthropic. Anthropic achieves ISO 42001 certification for responsible AI. https://www.anthropic.com/news/anthropic-achieves-iso-42001-certification-for-responsible-ai
An early frontier-lab ISO/IEC 42001 certification, described by the lab as responsible-AI governance rather than a guarantee that any particular model is safe.
[Anthropic 2025] Anthropic. Donating the model context protocol and establishing the agentic AI foundation. https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
Anthropic donates MCP to the Agentic AI Foundation, a Linux Foundation directed fund co-founded by Anthropic, Block, and OpenAI with support from Google, Microsoft, AWS, Cloudflare, and Bloomberg, to keep agentic AI standards neutral and community-driven.
[Anthropic 2025] Anthropic. The anthropic economic index. https://www.anthropic.com/economic-index
The Anthropic Economic Index maps millions of anonymized Claude conversations onto occupational tasks, showing usage concentrated in a narrow set of tasks and split between augmentation and automation.
[Anthropic 2026] Anthropic. Claude's new constitution. https://www.anthropic.com/news/claude-new-constitution
The January 2026 constitution is a book-length document written primarily for Claude itself, with a stated priority order of safety, ethics, Anthropic's guidelines, then helpfulness; Claude uses it directly in training to generate conversations, responses, and response rankings.
[Anthropic 2026] Anthropic. Context engineering: Memory, compaction, and tool clearing. https://platform.claude.com/cookbook/tool-use-context-engineering-context-engineering-tools
An Anthropic Cookbook guide comparing context engineering strategies (memory, compaction, and tool clearing) for long-running agents, covering when each applies, costs, and how they compose.
[Anthropic 2026] Anthropic. Automatic context compaction. https://platform.claude.com/cookbook/tool-use-automatic-context-compaction
Automatic context compaction monitors per-turn token usage and, past a configurable threshold, summarizes and replaces conversation history, cutting a five-ticket workflow from about 209K to 86K tokens.
[Anthropic 2026] Anthropic. Demystifying evals for AI agents. https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
Anthropic's practitioner guide explains how to design rigorous evals for AI agents, covering multi-turn task structures, grading strategies, and techniques that scale across real-world deployments.
[Anthropic 2026] Anthropic. Project Deal: our Claude-run marketplace experiment. https://www.anthropic.com/features/project-deal
An internal marketplace where employees' agents negotiated real purchases: agents on the stronger model bought cheaper and sold dearer on identical items, while participants rated their deals equally fair.
[Anthropic and Pattern Labs 2025] Anthropic and Pattern Labs. Confidential inference systems: Design principles and security risks. https://assets.anthropic.com/m/c52125297b85a42/original/Confidential_Inference_Paper.pdf
Design principles for confidential inference covering both directions of the trust problem: user data protected from the provider, and model weights protected from the infrastructure operator.
[anthropics/claude-code 2025] anthropics/claude-code. Massive token consumption: 1.67B tokens in 5 hours, issue #4095. https://github.com/anthropics/claude-code/issues/4095
A Claude Code bug report documenting 1.67B tokens consumed in 5 hours due to four confirmed failure modes: plan-execution loops, KV cache explosion, recursive hook calls, and API error retry loops.
[Apple 2023] Apple. MLX. https://github.com/ml-explore/mlx
MLX is an array framework for machine learning on Apple silicon, designed for efficient model training and inference on Apple hardware.
[Apple 2025] Apple. Apple intelligence foundation language models: Tech report 2025. arXiv preprint arXiv:2507.13575. https://arxiv.org/abs/2507.13575
Apple introduces a  3B-parameter on-device model with KV-cache sharing and 2-bit quantization-aware training, and a server model using a novel PT-MoE architecture combining track parallelism, MoE, and interleaved global-local attention for Apple Intelligence.
[Apple Machine Learning Research 2022] Apple Machine Learning Research. Deploying transformers on the apple neural engine. https://machinelearning.apple.com/research/neural-engine-transformers
Apple open-sources a PyTorch Transformer reference implementation optimized for the Apple Neural Engine, achieving up to 10x faster inference and 14x lower memory on-device versus the baseline.
[Apple Security Engineering and Architecture 2024] Apple Security Engineering and Architecture. Private Cloud Compute: A new frontier for AI privacy in the cloud. https://security.apple.com/blog/private-cloud-compute/
Apple's Private Cloud Compute (PCC) is a cloud AI inference system built on custom Apple silicon and a hardened OS that ensures personal user data processed by foundation models remains inaccessible even to Apple.
[ARC Prize Foundation 2026] ARC Prize Foundation. ARC-AGI-3: a new challenge for frontier agentic intelligence. arXiv preprint arXiv:2603.24621. https://arxiv.org/abs/2603.24621
ARC-AGI-3 moves the ARC family from static grids to hundreds of interactive turn-based environments with no instructions, where humans solve every level and frontier agents scored under 1
[ARC Prize Foundation 2026] ARC Prize Foundation. ARC-AGI leaderboard. https://arcprize.org/leaderboard
The live ARC-AGI leaderboard, showing ARC-AGI-2 top scores in the mid-eighties by June 2026, well past the average human score.
[Armilla 2025] Armilla. Armilla launches affirmative AI liability insurance backed by Lloyd's underwriters. https://www.armilla.ai/
An affirmative AI liability policy underwritten at Lloyd's whose trigger is AI underperformance, including hallucinations, model drift, and deviations from expected behavior.
[Arora et al. 2025] Arora, Wei, others. HealthBench: Evaluating large language models towards improved human health. arXiv preprint arXiv:2505.08775. https://arxiv.org/abs/2505.08775
HealthBench grades health conversations against 48,562 rubric criteria written by 262 physicians with practice experience across 60 countries, and meta-evaluates the model grader against physician grading before trusting its scores.
[Arriola et al. 2025] Arriola, Gokaslan, Chiu, Yang, Qi, Han, Sahoo, Kuleshov. Block diffusion: Interpolating between autoregressive and diffusion language models. https://arxiv.org/abs/2503.09573
Block Diffusion (BD3-LMs) interpolates between discrete diffusion and autoregressive models by applying diffusion within blocks of tokens, enabling variable-length generation and KV caching while achieving state-of-the-art perplexity among discrete diffusion models.
[Asai et al. 2023] Asai, Wu, Wang, Sil, Hajishirzi. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511. https://arxiv.org/abs/2310.11511
SELF-RAG trains a language model to adaptively retrieve passages on demand and critique its own generations using special reflection tokens, improving factuality without sacrificing versatility.
[Asai et al. 2023] Asai, Schick, Lewis, Chen, Izacard, Riedel, Hajishirzi, Yih. Task-aware retrieval with instructions. https://arxiv.org/abs/2211.09260
TART is a multi-task retrieval system trained on BERRI, a collection of  40 instruction-annotated datasets, that follows natural-language instructions to achieve state-of-the-art zero-shot retrieval on BEIR and LOTTE.
[Assran et al. 2025] Assran, Bardes, others. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985. https://arxiv.org/abs/2506.09985
V-JEPA 2 is a self-supervised video model pretrained on 1M hours of internet video that achieves state-of-the-art video understanding and enables zero-shot robot manipulation planning via a post-trained action-conditioned world model.
[Austin et al. 2021] Austin, Johnson, Ho, Tarlow, Berg. Structured denoising diffusion models in discrete state-spaces. https://arxiv.org/abs/2107.03006
D3PMs generalize denoising diffusion probabilistic models (DDPM) to discrete state spaces by replacing uniform corruption with structured transition matrices, achieving competitive text and image generation results.
[Austin et al. 2025] Austin, Douglas, Frostig, Levskaya, Chen, Vikram, Lebron, Choy, Ramasesh, Webson, Pope. How to scale your model. Google DeepMind. https://jax-ml.github.io/scaling-book/
A DeepMind/JAX systems book that explains LLM scaling on real TPU and GPU hardware, including rooflines, sharding, training, inference, serving, and profiling.
[Auth0 2026] Auth0. Identity for AI agents. https://auth0.com/blog/identity-for-ai-agents/
[Authors 2025] Authors. A survey on vision-language-action models: An action tokenization perspective. arXiv preprint arXiv:2507.01925. https://arxiv.org/abs/2507.01925
Survey of VLA models unified under an action tokenization framework, categorizing eight action token types and analyzing their trade-offs for embodied AI.
[Azar et al. 2023] Azar, Rowland, Piot, Guo, Calandriello, Valko, Munos. A general theoretical paradigm to understand learning from human preferences. arXiv preprint arXiv:2310.12036. https://arxiv.org/abs/2310.12036
This paper introduces ΨPO, a general preference optimization objective that unifies RLHF and DPO as special cases, and proposes IPO, which bypasses the Bradley-Terry assumption to avoid overfitting.
[Ba et al. 2016] Ba, Kiros, Hinton. Layer normalization. https://arxiv.org/abs/1607.06450
Layer normalization computes normalization statistics across all hidden units within a single layer and training case, eliminating batch-size constraints and stabilizing recurrent neural network training.
[Baevski et al. 2020] Baevski, Zhou, Mohamed, Auli. wav2vec 2.0: a framework for self-supervised learning of speech representations. https://arxiv.org/abs/2006.11477
wav2vec 2.0 learns speech representations via self-supervised masking and contrastive learning over quantized latents, enabling ASR with as little as ten minutes of labeled data.
[Bai et al. 2022] Bai, Kadavath, Kundu, Askell, Kernion, Jones, Chen, Goldie, Mirhoseini, McKinnon, Chen, Olsson, Olah, Hernandez, Drain, Ganguli, Li, Tran-Johnson, Perez, Kerr, Mueller, Ladish, Landau, Ndousse, Lukosuite, Lovitt, Sellitto, Elhage, Schiefer, Mercado, DasSarma, Lasenby, Larson, Ringer, Johnston, Kravec, El Showk, Fort, Lanham, Telleen-Lawton, Conerly, Henighan, Hume, Bowman, Hatfield-Dodds, Mann, Amodei, Joseph, McCandlish, Brown, Kaplan. Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073. https://arxiv.org/abs/2212.08073
Constitutional AI uses written principles, self-critique, revision, and AI feedback to train harmless but non-evasive assistant behavior.
[Bai et al. 2023] Bai, Bai, Chu, Cui, Dang, Deng, Fan, Ge, Han, Huang, others. Qwen technical report. https://arxiv.org/abs/2309.16609
Qwen is a family of large language models trained on up to 3 trillion tokens, covering base pretrained models, RLHF-aligned chat models, and specialized coding and mathematics variants.
[Bainbridge 1983] Bainbridge. Ironies of automation. Automatica. https://doi.org/10.1016/0005-1098(83)90046-8
Bainbridge names the paradox that automation removes routine practice while leaving humans responsible for rare, difficult interventions when automation fails.
[Bansal et al. 2021] Bansal, Wu, Zhou, Fok, Nushi, Kamar, Ribeiro, Weld. Does the whole exceed its parts? The effect of AI explanations on complementary team performance. https://arxiv.org/abs/2006.14779
Bansal et al. find that AI explanations did not improve complementary human-AI team performance and could increase acceptance of recommendations regardless of correctness.
[Barham and Isard 2019] Barham, Isard. Machine learning systems are stuck in a rut. https://dl.acm.org/doi/10.1145/3317550.3321441
[Baronio et al. 2025] Baronio, Marsella, Pan, Guo, Alberti. Kevin: Multi-turn RL for generating CUDA kernels. arXiv preprint arXiv:2507.11948. https://arxiv.org/abs/2507.11948
[Barres et al. 2025] Barres, Dong, Ray, Si, Narasimhan. τ²-Bench: Evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982. https://arxiv.org/abs/2506.07982
τ²-bench extends τ-bench to a dual-control telecom domain where both the agent and a simulated user act on a shared environment, separating reasoning errors from coordination failures.
[Baur and Strassen 1983] Baur, Strassen. The complexity of partial derivatives. Theoretical Computer Science. https://doi.org/10.1016/0304-3975(83)90110-X
[Baydin et al. 2018] Baydin, Pearlmutter, Radul, Siskind. Automatic differentiation in machine learning: a survey. Journal of Machine Learning Research. https://jmlr.org/papers/v18/17-468.html
The definitive survey of automatic differentiation for machine learning: forward and reverse modes, their cost asymmetry, and why AD is neither symbolic nor numerical differentiation.
[Becker et al. 2025] Becker, Rush, Barnes, Rein. Measuring the impact of early-2025 AI on experienced open-source developer productivity. https://arxiv.org/abs/2507.09089
In an RCT on 16 experienced open-source developers and 246 tasks, early-2025 AI tools slowed completion time by 19 percent despite participants expecting speedups.
[BehnamGhader et al. 2024] BehnamGhader, Adlakha, Mosbach, Bahdanau, Chapados, Reddy. LLM2Vec: Large language models are secretly powerful text encoders. https://arxiv.org/abs/2404.05961
LLM2Vec is an unsupervised three-step method (bidirectional attention, masked next-token prediction, SimCSE contrastive learning) that converts any decoder-only LLM into a strong text encoder, reaching state-of-the-art on MTEB.
[Benjamini and Hochberg 1995] Benjamini, Hochberg. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
Benjamini and Hochberg introduce false-discovery-rate control, a less conservative alternative to family-wise error control for settings with many simultaneous tests.
[Berdoz et al. 2026] Berdoz, Rugli, Wattenhofer. Can AI agents agree?. arXiv preprint arXiv:2603.01213. https://arxiv.org/abs/2603.01213
This paper evaluates LLM-based agents on a Byzantine consensus task and finds that valid agreement is unreliable even in benign settings, with failures dominated by liveness loss rather than value corruption.
[Bergstra et al. 2010] Bergstra, Breuleux, Bastien, Lamblin, Pascanu, Desjardins, Turian, Warde-Farley, Bengio. Theano: a CPU and GPU math compiler in python. https://proceedings.scipy.org/articles/Majora-92bf1922-003
The ancestor of the deep-learning framework: build a symbolic expression graph in Python, differentiate it as a graph transformation, and compile it to CPU or GPU code.
[Berkeley RDI 2025] Berkeley RDI. Trustworthy benchmarks for agents. https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/
Berkeley RDI researchers show that eight prominent AI agent benchmarks (SWE-bench, WebArena, OSWorld, GAIA, and others) can each be exploited to achieve near-perfect scores without solving any tasks.
[Besiroglu et al. 2024] Besiroglu, Erdil, Barnett, You. Chinchilla scaling: a replication attempt. arXiv preprint arXiv:2404.10102. https://arxiv.org/abs/2404.10102
Replicates the Chinchilla parametric fit and finds its third estimate inconsistent with the paper's own first two, with confidence intervals so narrow they would require hundreds of thousands of runs rather than the few hundred actually performed.
[Besta et al. 2023] Besta, Blach, Kubicek, Gerstenberger, Podstawski, Gianinazzi, Gajda, Lehmann, Niewiadomski, Nyczyk, Hoefler. Graph of thoughts: Solving elaborate problems with large language models. arXiv preprint arXiv:2308.09687. https://arxiv.org/abs/2308.09687
Graph of Thoughts generalizes chain and tree prompting into arbitrary graphs of LLM-generated thoughts, enabling aggregation, refinement, and feedback loops over intermediate reasoning units.
[Betker et al. 2023] Betker, Goh, Jing, Brooks, Wang, Li, Ouyang, Zhuang, Lee, Guo, Manassra, Dhariwal, Chu, Jiao, Ramesh. Improving image generation with better captions. https://cdn.openai.com/papers/dall-e-3.pdf
DALL-E 3 improves text-to-image prompt following by training on synthetic descriptive captions generated by a fine-tuned image captioner applied to the training dataset.
[Beurer-Kellner et al. 2025] Beurer-Kellner, Buesser, Creţu, Debenedetti, Dobos, Fabian, Fischer, Froelicher, Grosse, Naeff, Ozoani, Paverd, Tramèr, Volhejn. Design patterns for securing LLM agents against prompt injections. arXiv preprint arXiv:2506.08837. https://arxiv.org/abs/2506.08837
Researchers across ETH, Google, Microsoft, and Invariant systematize six design patterns that constrain an agent's structure so an injected instruction cannot redirect its privileged actions.
[Beyer et al. 2016] Beyer, Jones, Petoff, Murphy. Site reliability engineering: How google runs production systems. O'Reilly Media. https://sre.google/books/
Google's SRE book provides the operational vocabulary this chapter adapts to AI systems: service promises, error budgets, incident command, and learning from failure.
[Bhatt et al. 2025] Bhatt, Rushing, Kaufman, Tracy, Georgiev, Matolcsi, Khan, Shlegeris. Ctrl-z: Controlling AI agents via resampling. arXiv preprint arXiv:2504.10374. https://arxiv.org/abs/2504.10374
Ctrl-Z evaluates control protocols on BashBench's 257 multi-step sysadmin tasks and introduces resampling, rerunning suspicious actions for more evidence, cutting attack success from 58
[Biderman et al. 2023] Biderman, Schoelkopf, Anthony, Bradley, O'Brien, Hallahan, Khan, Purohit, Prashanth, Raff, Skowron, Sutawika, Wal. Pythia: a suite for analyzing large language models across training and scaling. Proceedings of the 40th International Conference on Machine Learning (ICML). https://arxiv.org/abs/2304.01373
Pythia is a suite of 16 LLMs from 70M to 12B parameters, each with 154 public checkpoints trained on the same data order, designed to study training dynamics and scaling.
[Bie et al. 2025] Bie, Cao, Chen, Du, others. LLaDA2.0: Scaling up diffusion language models to 100B. arXiv preprint arXiv:2512.15745. https://arxiv.org/abs/2512.15745
[Birgisson et al. 2014] Birgisson, Politz, Erlingsson, Taly, Vrable, Lentczner. Macaroons: Cookies with contextual caveats for decentralized authorization in the cloud. https://research.google/pubs/macaroons-cookies-with-contextual-caveats-for-decentralized-authorization-in-the-cloud/
Macaroons introduces bearer tokens with chained contextual caveats that enable decentralized, attenuated authorization for cloud services without a central authority.
[Black Forest Labs 2024] Black Forest Labs. Announcing black forest labs and the FLUX.1 model family. https://bfl.ai/announcements/24-08-01-bfl
The FLUX.1 flow-matching image model family, whose full-sampler and few-step distilled variants share one architecture but are priced per megapixel at an eightfold gap tracking their step counts.
[Bommasani et al. 2023] Bommasani, Klyman, Longpre, Kapoor, Maslej, Xiong, Zhang, Liang. The foundation model transparency index. https://arxiv.org/abs/2310.12941
The Foundation Model Transparency Index (FMTI) scores 10 major foundation model developers across 100 indicators covering upstream resources, model details, and downstream deployment practices.
[Borsos et al. 2023] Borsos, Marinier, Vincent, Kharitonov, Pietquin, Sharifi, Roblek, Teboul, Grangier, Tagliasacchi, Zeghidour. AudioLM: a language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing. https://arxiv.org/abs/2209.03143
AudioLM frames audio generation as language modeling over a hybrid of semantic tokens (from w2v-BERT) and acoustic tokens (from SoundStream), achieving both long-term coherence and high-quality synthesis for speech and piano continuation.
[Borsos et al. 2023] Borsos, Sharifi, Vincent, Kharitonov, Zeghidour, Tagliasacchi. SoundStorm: Efficient parallel audio generation. arXiv preprint arXiv:2305.09636. https://arxiv.org/abs/2305.09636
SoundStorm is a non-autoregressive audio generation model that uses bidirectional attention and confidence-based parallel decoding over RVQ token sequences, producing audio two orders of magnitude faster than AudioLM.
[Bourtoule et al. 2021] Bourtoule, Chandrasekaran, Choquette-Choo, Jia, Travers, Zhang, Lie, Papernot. Machine unlearning. https://arxiv.org/abs/1912.03817
This paper introduces SISA training, which partitions data into shards and slices to reduce the retraining cost of machine unlearning, achieving up to 4.63x speedup over full retraining.
[Bradford 2020] Bradford. The brussels effect: How the european union rules the world. Oxford University Press.
Bradford documents the Brussels Effect, how the EU exports its rules worldwide by setting de facto global standards through market access.
[Bradley and Terry 1952] Bradley, Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika. https://doi.org/10.2307/2334029
[Breck et al. 2017] Breck, Cai, Nielsen, Salib, Sculley. The ML test score: a rubric for ML production readiness and technical debt reduction. https://research.google/pubs/pub46555/
The ML Test Score provides a production-readiness rubric across data, model, infrastructure, and monitoring tests.
[Bricken et al. 2023] Bricken, Templeton, Batson, Chen, Jermyn, Conerly, Turner, Anil, Denison, Askell, Lasenby, Wu, Kravec, Schiefer, Maxwell, Joseph, Hatfield-Dodds, Tamkin, Nguyen, McLean, Burke, Hume, Carter, Henighan, Olah. Towards monosemanticity: Decomposing language models with dictionary learning. https://transformer-circuits.pub/2023/monosemantic-features/index.html
Sparse autoencoders applied via dictionary learning decompose a one-layer transformer's MLP activations into large numbers of interpretable, monosemantic features that overcome polysemanticity from superposition.
[Brill 2024] Brill. Neural scaling laws rooted in the data distribution. arXiv preprint arXiv:2412.07942. https://arxiv.org/abs/2412.07942
Uses percolation theory on natural data to derive both the discrete-quanta and the data-manifold pictures of scaling from first principles, unifying the two explanations as regimes on either side of a percolation threshold.
[Brohan et al. 2023] Brohan, Brown, others. RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818. https://arxiv.org/abs/2307.15818
RT-2 co-fine-tunes large VLMs on robotic trajectory data by encoding robot actions as text tokens, producing VLA models that generalize to novel objects and exhibit emergent semantic reasoning.
[Brooks et al. 2024] Brooks, Peebles, others. Video generation models as world simulators. https://openai.com/index/video-generation-models-as-world-simulators/
The Sora technical report: a diffusion transformer over spatiotemporal patches of video latents, with compute as the axis along which sample quality scales, up to a minute of generated video.
[Brown et al. 2020] Brown, Mann, Ryder, Subbiah, Kaplan, Dhariwal, Neelakantan, Shyam, Sastry, Askell, Agarwal, Herbert-Voss, Krueger, Henighan, Child, Ramesh, Ziegler, Wu, Winter, Hesse, Chen, Sigler, Litwin, Gray, Chess, Clark, Berner, McCandlish, Radford, Sutskever, Amodei. Language models are few-shot learners. arXiv preprint arXiv:2005.14165. https://arxiv.org/abs/2005.14165
GPT-3, a 175-billion-parameter autoregressive language model, achieves strong few-shot performance on NLP benchmarks without gradient updates or fine-tuning, using only in-context demonstrations.
[Brown et al. 2024] Brown, Juravsky, Ehrlich, Clark, Le, Ré, Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787. https://arxiv.org/abs/2407.21787
This paper shows that repeatedly sampling from LLMs scales coverage log-linearly over four orders of magnitude, enabling weaker models to exceed single-sample state-of-the-art on coding and proof benchmarks.
[Brown-Cohen et al. 2023] Brown-Cohen, Irving, Piliouras. Scalable AI safety via doubly-efficient debate. arXiv preprint arXiv:2311.14125. https://arxiv.org/abs/2311.14125
This paper introduces doubly-efficient debate protocols where an honest AI prover can verify polynomial-time computations, including stochastic ones, using only a constant number of human queries, reducing the exponential simulation cost of prior debate frameworks.
[Brynjolfsson et al. 2023] Brynjolfsson, Li, Raymond. Generative AI at work. https://arxiv.org/abs/2304.11771
A staggered rollout of a conversational AI assistant raised customer-support productivity by 15 percent on average, with the largest gains for less experienced workers.
[Buçinca et al. 2021] Buçinca, Malaya, Gajos. To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. https://arxiv.org/abs/2102.09692
Buçinca, Malaya, and Gajos show that cognitive forcing interventions reduce overreliance on AI recommendations more than simple explanation displays, with usability trade-offs.
[Burges 2010] Burges. From RankNet to LambdaRank to LambdaMART: An overview. https://www.microsoft.com/en-us/research/publication/from-ranknet-to-lambdarank-to-lambdamart-an-overview/
[Burns et al. 2024] Burns, Izmailov, Kirchner, Baker, Gao, Aschenbrenner, Chen, Ecoffet, Joglekar, Leike, Sutskever, Wu. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390. https://arxiv.org/abs/2312.09390
Weak-to-strong generalization shows that weak labels can elicit some stronger-model capability, but naive fine-tuning remains far from recovering the full strong model.
[ByteDance Seed 2025] ByteDance Seed. Robust LLM training infrastructure at ByteDance (ByteRobust). arXiv preprint arXiv:2509.16293. https://arxiv.org/abs/2509.16293
ByteRobust is ByteDance's production LLM training infrastructure system for fault detection and recovery at scale, achieving 97
[Cai et al. 2024] Cai, Li, Geng, Peng, Lee, Chen, Dao. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. https://arxiv.org/abs/2401.10774
MEDUSA accelerates LLM inference by adding multiple extra decoding heads plus tree-based attention to predict and verify several tokens per step, achieving 2.2-2.8x speedup without a separate draft model.
[California State Legislature 2025] California State Legislature. SB 53: Transparency in frontier artificial intelligence act. https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202520260SB53
[Campbell et al. 2020] Campbell, Bradley, Tschofenig. RFC 8707: Resource indicators for OAuth 2.0. https://datatracker.ietf.org/doc/html/rfc8707
RFC 8707 defines an OAuth 2.0 extension adding request parameters that let a client explicitly indicate to an authorization server which protected resource(s) it is requesting access to.
[Cardwell 2025] Cardwell. GE vernova expects to end 2025 with an 80-GW gas turbine backlog that stretches into 2029. https://www.utilitydive.com/news/ge-vernova-gas-turbine-investor/807662/
GE Vernova expects to close 2025 with an 80-GW gas turbine backlog stretching to 2029, with reservations projected to sell out through 2030 by end of 2026.
[Carlini et al. 2021] Carlini, Tramèr, Wallace, Jagielski, Herbert-Voss, Lee, Roberts, Brown, Song, Erlingsson, Oprea, Raffel. Extracting training data from large language models. https://arxiv.org/abs/2012.07805
Carlini et al. show that large language models memorize training data, and that black-box query access to GPT-2 can extract verbatim text including personally identifiable information.
[Carlini et al. 2023] Carlini, Ippolito, Jagielski, Lee, Tramer, Zhang. Quantifying memorization across neural language models. https://arxiv.org/abs/2202.07646
Carlini et al. quantify three log-linear relationships showing that LLM memorization of training data grows with model capacity, data duplication, and context length, and is more prevalent than previously believed.
[Carlini et al. 2023] Carlini, Jagielski, Choquette-Choo, Paleka, Pearce, Anderson, Terzis, Thomas, Tramèr. Poisoning web-scale training datasets is practical. arXiv preprint arXiv:2302.10149. https://arxiv.org/abs/2302.10149
Two practical poisoning attacks on web-scale datasets: buying expired domains the crawler still trusts, and editing pages just before a snapshot, at a cost of about sixty dollars for a major dataset.
[CBS News 2024] CBS News. Google strikes $60 million deal with reddit, allowing search giant to train AI models on human posts. https://www.cbsnews.com/news/google-reddit-60-million-deal-ai-training/
[Cemri et al. 2025] Cemri, Pan, Yang, Agrawal, Chopra, Tiwari, Keutzer, Parameswaran, Klein, Ramchandran, Zaharia, Gonzalez, Stoica. Why do multi-agent LLM systems fail?. https://openreview.net/forum?id=fAjbYBmonr
MAST introduces a 14-mode failure taxonomy for multi-agent LLM systems, paired with MAST-Data (1,642 annotated traces across 7 frameworks) and an LLM-as-judge annotation pipeline to diagnose why MAS fail.
[Center for AI Safety and Scale AI 2026] Center for AI Safety and Scale AI. Humanity's last exam. https://agi.safe.ai/
Humanity's Last Exam (HLE) is a 2,500-question benchmark of expert-level problems across diverse fields, designed to be among the hardest tests for frontier AI models.
[Cerebras Systems 2024] Cerebras Systems. Cerebras announces third-generation wafer-scale engine (WSE-3). https://www.cerebras.ai/press-release/cerebras-announces-third-generation-wafer-scale-engine
Cerebras announces WSE-3, a 5nm wafer-scale chip with 4 trillion transistors and 900,000 cores delivering 125 petaflops, doubling WSE-2 performance at the same power and price.
[Chameleon Team 2024] Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818. https://arxiv.org/abs/2405.09818
Chameleon is a family of early-fusion token-based mixed-modal foundation models that understand and generate arbitrarily interleaved image and text sequences using a single unified transformer trained with QK-norm for stability.
[Chan et al. 2016] Chan, Jaitly, Le, Vinyals. Listen, attend and spell. https://arxiv.org/abs/1508.01211
LAS is an end-to-end sequence-to-sequence speech recognizer with a pyramidal RNN encoder and attention-based character decoder, achieving 14.1
[Chan et al. 2020] Chan, Saharia, Hinton, Norouzi, Jaitly. Imputer: Sequence modelling via imputation and dynamic programming. https://arxiv.org/abs/2002.08926
Imputer is an iterative non-autoregressive sequence model for speech recognition that uses dynamic programming to marginalize over alignments and generation orders, achieving constant-step decoding and outperforming CTC on LibriSpeech.
[Chaney et al. 2018] Chaney, Stewart, Engelhardt. How algorithmic confounding in recommendation systems increases homogeneity and decreases utility. https://arxiv.org/abs/1710.11214
[Chanin et al. 2024] Chanin, Wilken-Smith, Dulka, Bhatnagar, Golechha, Bloom. A is for absorption: Studying feature splitting and absorption in sparse autoencoders. arXiv preprint arXiv:2409.14507. https://arxiv.org/abs/2409.14507
This paper identifies "feature absorption" in sparse autoencoders (SAEs), where hierarchical features cause SAE latents to silently fail to activate on tokens they should track, undermining reliable LLM interpretability.
[Chen and Guestrin 2016] Chen, Guestrin. XGBoost: a scalable tree boosting system. https://arxiv.org/abs/1603.02754
The systems design of gradient-boosted trees at scale: cache-aware layout, sparsity handling, and out-of-core computation, with the paper's own count that 17 of 29 published Kaggle winning solutions in 2015 used it.
[Chen et al. 2016] Chen, Xu, Zhang, Guestrin. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174. https://arxiv.org/abs/1604.06174
Keep activations at only O(sqrt(n)) of a network's layers and recompute the rest during the backward pass: a 1000-layer network's training memory drops from 48 GB to 7 GB for about 30
[Chen et al. 2018] Chen, Moreau, Jiang, Zheng, Yan, Shen, Cowan, Wang, Hu, Ceze, Guestrin, Krishnamurthy. TVM: An automated end-to-end optimizing compiler for deep learning. https://www.usenix.org/conference/osdi18/presentation/chen
TVM is an end-to-end optimizing compiler for deep learning that automatically generates hardware-specific code for CPUs, GPUs, FPGAs, and ASICs without manual operator tuning.
[Chen et al. 2021] Chen, Tworek, Jun, Yuan, Pinto, Kaplan, Edwards, Burda, Joseph, Brockman, Ray, Puri, Krueger, Petrov, Khlaaf, Sastry, Mishkin, Chan, Gray, Ryder, Pavlov, Power, Kaiser, Bavarian, Winter, Tillet, Such, Cummings, Plappert, Chantzis, Barnes, Herbert-Voss, Guss, Nichol, Paino, Tezak, Tang, Babuschkin, Balaji, Jain, Saunders, Hesse, Carr, Leike, Achiam, Misra, Morikawa, Radford, Knight, Brundage, Murati, Mayer, Welinder, McGrew, Amodei, McCandlish, Sutskever, Zaremba. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. https://arxiv.org/abs/2107.03374
This paper introduces Codex, a GPT model fine-tuned on GitHub code, and releases HumanEval, a 164-problem benchmark measuring functional correctness via pass@k unit-test evaluation.
[Chen et al. 2022] Chen, Wang, Chen, Wu, Liu, Chen, Li, Kanda, Yoshioka, Xiao, others. WavLM: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing. https://arxiv.org/abs/2110.13900
WavLM is a self-supervised speech pre-training model that jointly learns masked speech prediction and denoising to achieve state-of-the-art performance across full-stack speech tasks including ASR, speaker verification, separation, and diarization.
[Chen et al. 2022] Chen, Ma, Wang, Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588. https://arxiv.org/abs/2211.12588
Program of Thoughts prompting asks LLMs to express numerical reasoning as executable programs, separating reasoning decomposition from exact computation.
[Chen et al. 2023] Chen, Wong, Chen, Tian. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595. https://arxiv.org/abs/2306.15595
Extends RoPE-based LLM context windows by linearly down-scaling position indices, avoiding unstable extrapolation and preserving original-window quality after limited fine-tuning.
[Chen et al. 2023] Chen, Borgeaud, Irving, Lespiau, Sifre, Jumper. Accelerating large language model decoding with speculative sampling. https://arxiv.org/abs/2302.01318
Speculative sampling uses a small draft model to propose tokens that a large target model verifies in parallel, achieving 2–2.5x decoding speedup on Chinchilla without altering the output distribution.
[Chen et al. 2024] Chen, Liu, Zhou, Liu, Tan, Li, Zhao, Qian, Wei. VALL-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers. arXiv preprint arXiv:2406.05370. https://arxiv.org/abs/2406.05370
VALL-E 2 is a neural codec language model for zero-shot TTS that introduces repetition aware sampling and grouped code modeling to achieve human parity on LibriSpeech and VCTK.
[Chen et al. 2024] Chen, Niu, Ma, Deng, Wang, Zhao, Yu, Chen. F5-TTS: a fairytaler that fakes fluent and faithful speech with flow matching. arXiv preprint arXiv:2410.06885. https://arxiv.org/abs/2410.06885
F5-TTS is a non-autoregressive TTS system using flow matching with DiT and ConvNeXt, achieving zero-shot voice cloning without phoneme alignment or duration models.
[Chen et al. 2024] Chen, Wu, Wang, Su, Chen, Xing, Zhong, Zhang, Zhu, Lu, Li, Luo, Lu, Qiao, Dai. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/2312.14238
InternVL scales a vision transformer encoder to 6 billion parameters and progressively aligns it with LLMs via a language middleware, achieving state-of-the-art results across 32 visual-linguistic benchmarks.
[Chen et al. 2024] Chen, Zhao, Liu, Bai, Lin, Zhou, Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models. https://arxiv.org/abs/2403.06764
FastV prunes visual tokens in large vision-language models after layer 2 based on attention scores, achieving 45
[Chen et al. 2024] Chen, Xiang, Xiao, Song, Li. AgentPoison: Red-teaming LLM agents via poisoning memory or knowledge bases. arXiv preprint arXiv:2407.12784. https://arxiv.org/abs/2407.12784
[Chen et al. 2025] Chen, Hao, Liu, Huang, Zeng, Yu, Li, Wang, Gan, Huang, Liu, Wang, Lian, Yin, Wang, Liu. ACEBench: a comprehensive evaluation of LLM tool usage. https://aclanthology.org/2025.findings-emnlp.697.pdf
ACEBench is a comprehensive benchmark for evaluating LLM tool usage across normal, special (imperfect instructions), and agent (multi-turn) scenarios with LLM-free automated assessment.
[Cheng et al. 2024] Cheng, Ozga, Valdez, Ahmed, Gu, Jamjoom, Franke, Bottomley. Intel TDX demystified: a top-down approach. ACM Computing Surveys. https://dl.acm.org/doi/10.1145/3652597
A top-down academic treatment of Intel's trust-domain VMs: the architecture, the attestation flow, and the trust boundaries, without requiring the vendor specification.
[Chervonyi et al. 2025] Chervonyi, Trinh, Olsak, Yang, Nguyen, Menegali, Jung, Kim, Verma, Le, Luong. Gold-medalist performance in solving olympiad geometry with AlphaGeometry2. arXiv preprint arXiv:2502.03544. https://arxiv.org/abs/2502.03544
AlphaGeometry2 improves language coverage and symbolic search for Olympiad geometry, showing how neural proposal and symbolic verification can work together.
[Chetlur et al. 2014] Chetlur, Woolley, Vandermersch, Cohen, Tran, Catanzaro, Shelhamer. cuDNN: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759. https://arxiv.org/abs/1410.0759
The library that welded every 2010s framework to NVIDIA silicon: optimized deep-learning primitives beneath Caffe, Torch, and their successors, at the moment the field industrialized.
[Chhabria 2025] Chhabria. Kadrey v. meta platforms, inc., order on summary judgment, no. 3:23-cv-03417 (N.D. cal.). https://law.justia.com/cases/federal/district-courts/california/candce/3:2023cv03417/415175/598/
Judge Chhabria grants Meta summary judgment on fair use for LLM training, while noting in dicta that a developed market-dilution record might have carried the authors' case.
[Chhikara et al. 2025] Chhikara, Khant, Aryan, Singh, Yadav. Mem0: Building production-ready AI agents with scalable long-term memory. https://arxiv.org/abs/2504.19413
Mem0 is a scalable memory architecture for AI agents that dynamically extracts, consolidates, and retrieves salient facts across sessions, achieving 26
[Chhikara et al. 2025] Chhikara, Khant, Aryan, Singh, Yadav. Mem0: Building production-ready AI agents with scalable long-term memory. arXiv preprint arXiv:2504.19413. https://arxiv.org/abs/2504.19413
The memory-as-a-service pipeline: extract, consolidate, retrieve, with large reported latency and token savings over full-context, whose own tables show the full-context baseline scoring higher on quality.
[Chiang et al. 2024] Chiang, Zheng, Sheng, Angelopoulos, Li, Li, Zhang, Zhu, Jordan, Gonzalez, Stoica. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132. https://arxiv.org/abs/2403.04132
Chatbot Arena is an open crowdsourced platform that evaluates LLMs via pairwise human preference votes, using Bradley-Terry ranking over 240K crowd-collected comparisons.
[Chollet et al. 2025] Chollet, Knoop, others. ARC-AGI-2: a new challenge for frontier AI reasoning systems. arXiv preprint arXiv:2505.11831. https://arxiv.org/abs/2505.11831
ARC-AGI-2 is an upgraded abstract reasoning benchmark with harder, less brute-forcible tasks and large-scale first-party human baselines, designed to better measure progress toward general AI beyond ARC-AGI-1.
[Chowdhery et al. 2022] Chowdhery, Narang, Devlin, Bosma, Mishra, Roberts, Barham, Chung, Sutton, Gehrmann, others. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311. https://arxiv.org/abs/2204.02311
Trains PaLM, a 540B-parameter dense Transformer, on 6144 TPU v4 chips with the Pathways system, to study how scale improves few-shot performance across many tasks.
[Christiano et al. 2017] Christiano, Leike, Brown, Martic, Legg, Amodei. Deep reinforcement learning from human preferences. https://arxiv.org/abs/1706.03741
This paper shows that deep RL agents can learn complex behaviors from non-expert human preferences over trajectory segment pairs, requiring feedback on less than 1
[Christiano et al. 2018] Christiano, Shlegeris, Amodei. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575. https://arxiv.org/abs/1810.08575
Iterated amplification trains on difficult tasks by recursively decomposing them into easier subproblems that weak experts can help answer.
[Chroma 2025] Chroma. Chroma: Open-source embedding database. https://www.trychroma.com/
Chroma is an open-source AI search infrastructure providing vector, full-text, regex, and metadata search built on object storage, available under Apache 2.0.
[Chung et al. 2021] Chung, Zhang, Han, Chiu, Qin, Pang, Wu. w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. https://arxiv.org/abs/2108.06209
w2v-BERT combines wav2vec 2.0 contrastive learning with masked language modeling into a single end-to-end self-supervised framework for speech representation learning, achieving competitive ASR results on LibriSpeech.
[Clark et al. 2021] Clark, August, Serrano, Haduong, Gururangan, Smith. All that's 'human' is not gold: Evaluating human evaluation of generated text. arXiv preprint arXiv:2107.00061. https://arxiv.org/abs/2107.00061
Clark et al. show that untrained evaluators often identify GPT-3 generated text only at chance level, motivating evaluator training and clearer protocols.
[Cloudflare 2025] Cloudflare. Introducing pay per crawl: enabling content owners to charge AI crawlers for access. https://blog.cloudflare.com/introducing-pay-per-crawl/
The first at-scale HTTP 402 deployment: publishers set a per-request price for AI crawlers, who negotiate and pay through headers, per-fetch price discovery for machine readers of the web.
[Cloudflare 2025] Cloudflare. The age of agents: cryptographically recognizing agent traffic. https://blog.cloudflare.com/signed-agents/
Signed agents extend Cloudflare's verified-bot program to user-controlled agents: RFC 9421 HTTP message signatures against published keys, the emerging registry of who a robot is.
[CNBC 2025] CNBC. Michael Burry accuses AI hyperscalers of artificially boosting earnings. https://www.cnbc.com/2025/11/11/big-short-investor-michael-burry-accuses-ai-hyperscalers-of-artificially-boosting-earnings.html
Reports Michael Burry's claim that hyperscalers will understate depreciation by roughly $176B over 2026-2028 by stretching accelerator useful lives past their two-to-three-year economic life.
[CNBC 2025] CNBC. AI GPU depreciation: CoreWeave, nvidia respond to michael burry. https://www.cnbc.com/2025/11/14/ai-gpu-depreciation-coreweave-nvidia-michael-burry.html
Nvidia and CoreWeave respond that prior-generation accelerators like the A100 remain highly utilized and rebook near original pricing, defending four-to-six-year depreciation schedules.
[CNBC 2026] CNBC. Amazon wins court order to block Perplexity's AI shopping agent. https://www.cnbc.com/2026/03/10/amazon-wins-court-order-to-block-perplexitys-ai-shopping-agent.html
[Coalition for Content Provenance and Authenticity 2024] Coalition for Content Provenance and Authenticity. C2PA technical specification. https://c2pa.org/specifications/
[Cobbe et al. 2021] Cobbe, Kosaraju, Bavarian, Chen, Jun, Kaiser, Plappert, Tworek, Hilton, Nakano, Hesse, Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. https://arxiv.org/abs/2110.14168
Training Verifiers introduces GSM8K and shows that sampling many solutions then selecting with a verifier can outperform directly fine-tuning the generator on math word problems.
[Cohen 1960] Cohen. A coefficient of agreement for nominal scales. Educational and Psychological Measurement. https://doi.org/10.1177/001316446002000104
Cohen's kappa measures agreement between two raters after adjusting for agreement expected by chance, making raw percent agreement harder to misuse.
[Coinbase 2025] Coinbase. Introducing x402: a new standard for internet-native payments. https://www.coinbase.com/developer-platform/discover/launches/x402
Revives HTTP 402 as a concrete protocol: a server prices the request, the client retries with a signed stablecoin payment in a header, and a facilitator settles on-chain.
[Collobert et al. 2011] Collobert, Kavukcuoglu, Farabet. Torch7: a matlab-like environment for machine learning. https://ronan.collobert.com/pub/matos/2011_torch7_nipsw.pdf
[Computer Weekly 2023] Computer Weekly. Google saves almost $3bn by running servers for six years. https://www.computerweekly.com/news/366557152/Google-saves-almost-3bn-by-running-servers-for-six-years
Reports how Alphabet's 2023 extension of server useful life from four to six years reduced depreciation by $3.9B and lifted net income by $3.0B, part of an industry-wide pattern.
[Construct Validity Review 2025] Construct Validity Review. Measuring what matters: Construct validity in large language model benchmarks. https://openreview.net/forum?id=mdA5lVvNcU
[CoreWeave 2023] CoreWeave. CoreWeave secures $2.3 billion debt financing facility. https://www.prnewswire.com/news-releases/coreweave-secures-2-3-billion-debt-financing-facility-led-by-magnetar-capital-and-blackstone-301892706.html
The first large GPU-collateralized debt facility, lending against contracted accelerator revenue rather than the chips' resale value.
[Costan and Devadas 2016] Costan, Devadas. Intel SGX explained. https://eprint.iacr.org/2016/086
The definitive explainer of Intel SGX: how enclaves, measurement, and attestation actually work at the silicon level, still the best single on-ramp to trusted execution.
[Costanza-Chock et al. 2022] Costanza-Chock, Raji, Buolamwini. Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem. https://dl.acm.org/doi/10.1145/3531146.3533213
A field scan finding that algorithmic audit has no settled meaning, and that unverifiable audit claims can exacerbate rather than mitigate harm, the source for the audit-washing critique.
[Cottier et al. 2024] Cottier, Rahman, Fattorini, Maslej, Owen. How much does it cost to train frontier AI models?. https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models
Epoch AI analysis finds frontier AI model training costs have grown 2-3x annually for eight years, projecting the largest models could exceed one billion dollars by 2027.
[Cottier et al. 2024] Cottier, Rahman, Fattorini, Maslej, Owen. The rising costs of training frontier AI models. https://arxiv.org/abs/2405.21015
Cottier et al. build a detailed cost model showing frontier AI training costs have grown 2.4x per year since 2016, with GPT-4 and Gemini Ultra each costing tens of millions of dollars.
[Council of the European Union 2026] Council of the European Union. Artificial intelligence: Council and Parliament agree to simplify and streamline rules. https://www.consilium.europa.eu/en/press/press-releases/2026/05/07/artificial-intelligence-council-and-parliament-agree-to-simplify-and-streamline-rules/
[Council of the European Union 2026] Council of the European Union. Artificial intelligence: Council and parliament agree to simplify and streamline rules. https://www.consilium.europa.eu/en/press/press-releases/2026/05/07/artificial-intelligence-council-and-parliament-agree-to-simplify-and-streamline-rules/
[Council of the European Union 2026] Council of the European Union. Artificial intelligence: Council gives final green light to simplify and streamline rules. https://www.consilium.europa.eu/en/press/press-releases/2026/06/29/artificial-intelligence-council-gives-final-green-light-to-simplify-and-streamline-rules/
[Council on Foreign Relations 2025] Council on Foreign Relations. China's AI chip deficit: Why huawei can't catch nvidia and U.S. export controls should remain. https://www.cfr.org/articles/chinas-ai-chip-deficit-why-huawei-cant-catch-nvidia-and-us-export-controls-should-remain
A CFR policy analysis argues Huawei cannot match Nvidia in AI chip performance and that U.S. export controls on GPU exports to China should be maintained.
[Covington et al. 2016] Covington, Adams, Sargin. Deep neural networks for YouTube recommendations. https://dl.acm.org/doi/10.1145/2959100.2959190
The paper that fixed the industry's two-stage vocabulary: a candidate-generation network retrieves hundreds of items from millions, and a ranking network scores the survivors, the direct structural ancestor of retrieve-then-rerank in RAG.
[Cui et al. 2023] Cui, Yuan, Ding, Yao, He, Zhu, Ni, Xie, Xie, Lin, Liu, Sun. UltraFeedback: Boosting language models with scaled AI feedback. arXiv preprint arXiv:2310.01377. https://arxiv.org/abs/2310.01377
UltraFeedback uses scaled GPT-4 feedback over 250,000 conversations, emphasizing diversity and bias mitigation as the factors that make AI feedback useful for alignment.
[Cui et al. 2025] Cui, Fu, Zhang, Wang, Zuo. Free-MAD: Consensus-free multi-agent debate. arXiv preprint arXiv:2509.11035. https://arxiv.org/abs/2509.11035
FREE-MAD is a consensus-free multi-agent debate framework for LLMs that replaces majority voting with a score-based decision mechanism evaluating full debate trajectories, reducing token cost to a single round.
[Cui et al. 2025] Cui, Chiang, Stoica, Hsieh. OR-bench: An over-refusal benchmark for large language models. https://arxiv.org/abs/2405.20947
OR-Bench is the first large-scale over-refusal benchmark for LLMs, comprising 80,000 automatically generated safe prompts across 10 categories to measure and compare over-refusal rates across 32 models.
[Cunningham et al. 2023] Cunningham, Ewart, Riggs, Huben, Sharkey. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600. https://arxiv.org/abs/2309.08600
Sparse autoencoders trained on language model internal activations recover monosemantic, interpretable features by decomposing polysemantic neurons via superposition into a sparse overcomplete dictionary.
[Cursor 2025] Cursor. Checkpoints in the agent workflow. https://stevekinney.com/courses/ai-development/cursor-checkpoints
Cursor automatically snapshots codebase state before every AI Agent edit, storing ephemeral local checkpoints outside Git so users can roll back unwanted changes.
[D'Amour et al. 2020] D'Amour, Heller, Moldovan, Adlam, Alipanahi, Beutel, Chen, Deaton, Eisenstein, Hoffman, others. Underspecification presents challenges for credibility in modern machine learning. arXiv preprint arXiv:2011.03395. https://arxiv.org/abs/2011.03395
This paper argues that modern ML pipelines are often underspecified: many predictors have similar held-out performance but behave differently in deployment.
[Dabney et al. 2020] Dabney, Kurth-Nelson, Uchida, Starkweather, Hassabis, Munos, Botvinick. A distributional code for value in dopamine-based reinforcement learning. Nature.
Dabney et al. find that dopamine neurons represent a distribution over future rewards, giving biological evidence for distributional reinforcement learning.
[Dai et al. 2024] Dai, Deng, Zhao, Xu, Gao, Chen, Li, Zeng, Yu, Wu, Xie, Li, Huang, Luo, Ruan, Sui, Liang. DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models. https://arxiv.org/abs/2401.06066
DeepSeekMoE proposes fine-grained expert segmentation and shared expert isolation in MoE language models to achieve stronger expert specialization, matching dense model performance with far less computation.
[Dao et al. 2022] Dao, Fu, Ermon, Rudra, Ré. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. https://arxiv.org/abs/2205.14135
FlashAttention is an IO-aware exact attention algorithm that uses tiling and recomputation to minimize HBM accesses, achieving faster wall-clock training and linear memory usage in sequence length.
[Dao 2023] Dao. FlashAttention-2: Faster attention with better parallelism and work partitioning. https://arxiv.org/abs/2307.08691
FlashAttention-2 improves GPU work partitioning and parallelism across sequence length to achieve roughly 2x speedup over FlashAttention, reaching 50-73
[Dao and Gu 2024] Dao, Gu. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality. https://arxiv.org/abs/2405.21060
Mamba-2 unifies selective SSMs and attention via a structured state space duality (SSD) framework over semiseparable matrices, yielding a 2-8x faster SSM layer competitive with Transformers on language modeling.
[Dathathri et al. 2024] Dathathri, See, Ghaisas, Huang, McAdam, Welbl, Bachani, Kaskasoli, Stanforth, Matejovicova, Hayes, Vyas, Merey, Brown-Cohen, Bunel, Balle, Cemgil, Ahmed, Stacpoole, Shumailov, Baetu, Gowal, Hassabis, Kohli. Scalable watermarking for identifying large language model outputs. Nature. https://www.nature.com/articles/s41586-024-08025-4
SynthID-Text productionizes generation-time watermarking with negligible latency, was deployed on live Gemini traffic in a 20-million-response test, and is open-sourced.
[Davis 2025] Davis. Durable execution meets AI. https://temporal.io/blog/durable-execution-meets-ai-why-temporal-is-the-perfect-foundation-for-ai
Temporal's Durable Execution model satisfies every core reliability requirement of LLM-powered AI agent applications because agents are distributed systems that need resilience against transient failures.
[Dean et al. 2012] Dean, Corrado, Monga, Chen, Devin, Mao, Ranzato, Senior, Tucker, Yang, Le, Ng. Large scale distributed deep networks. https://proceedings.neurips.cc/paper/2012/hash/6aca97005c68f1206823815f66102863-Abstract.html
[Dean and Barroso 2013] Dean, Barroso. The tail at scale. ACM. https://research.google/pubs/the-tail-at-scale/
Dean and Barroso explain why tail latency dominates large-scale distributed systems and present hedging and redundancy techniques to reduce high-percentile response times.
[Debenedetti et al. 2024] Debenedetti, Zhang, Balunović, Beurer-Kellner, Fischer, Tramèr. AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. https://arxiv.org/abs/2406.13352
AgentDojo is an extensible benchmark framework with 97 realistic tasks and 629 security test cases for evaluating LLM agent robustness against prompt injection attacks.
[Debenedetti et al. 2025] Debenedetti, Shumailov, Fan, Hayes, Carlini, Fabian, Kern, Shi, Terzis, Tramèr. Defeating prompt injections by design. arXiv preprint arXiv:2503.18813. https://arxiv.org/abs/2503.18813
CaMeL is a system-layer defense against prompt injection in LLM agents that enforces capability-based security policies on control and data flows without modifying the underlying model, solving 77
[DeepSeek-AI 2024] DeepSeek-AI. DeepSeek-V3 technical report. https://arxiv.org/abs/2412.19437
Reports DeepSeek-V3, a 671B-parameter Mixture-of-Experts model with 37B active per token, trained on 14.8T tokens with fp8 matmuls and auxiliary-loss-free load balancing, rivaling closed models at low cost.
[DeepSeek-AI 2024] DeepSeek-AI. DeepSeek-V2: a strong, economical, and efficient mixture-of-experts language model. https://arxiv.org/abs/2405.04434
DeepSeek-V2 is a 236B MoE language model using Multi-head Latent Attention (MLA) and DeepSeekMoE to reduce KV cache by 93.3
[DeepSeek-AI 2025] DeepSeek-AI. DeepSeek-V3.2: Pushing the frontier of open large language models. arXiv preprint arXiv:2512.02556. https://arxiv.org/abs/2512.02556
Introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention driven by a lightning indexer, first shipped in DeepSeek-V3.2-Exp, cutting long-context attention cost from O(L^2) toward O(Lk) at near-identical output quality.
[DeepSeek-AI 2026] DeepSeek-AI. DeepSeek-V4-pro. https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
[Défossez et al. 2023] Défossez, Copet, Synnaeve, Adi. High fidelity neural audio compression. Transactions on Machine Learning Research (TMLR). https://arxiv.org/abs/2210.13438
EnCodec is a real-time neural audio codec using a streaming encoder-decoder with residual vector quantization and adversarial training, achieving state-of-the-art compression at 1.5 to 24 kbps.
[Défossez et al. 2024] Défossez, Mazaré, Orsini, Royer, Pérez, Jégou, Grave, Zeghidour. Moshi: a speech-text foundation model for real-time dialogue. arXiv preprint arXiv:2410.00037. https://arxiv.org/abs/2410.00037
Moshi is a real-time full-duplex speech-text foundation model that generates speech directly from speech using parallel audio streams and an Inner Monologue method, achieving 160ms theoretical latency.
[Delétang et al. 2024] Delétang, Ruoss, Duquenne, Catt, Genewein, Veness, Hutter, others. Language modeling is compression. https://arxiv.org/abs/2309.10668
Large language models are powerful general-purpose lossless compressors: Chinchilla 70B compresses ImageNet patches to 43.4
[Dell'Acqua et al. 2023] Dell'Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon, Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. https://www.hbs.edu/faculty/Pages/item.aspx?num=64700
The HBS field experiment finds large gains for tasks inside the AI capability frontier but worse accuracy on a task outside it, motivating the jagged-frontier framing.
[Dettmers et al. 2022] Dettmers, Lewis, Belkada, Zettlemoyer. llm.int8(): 8-bit matrix multiplication for transformers at scale. https://arxiv.org/abs/2208.07339
LLM.int8() halves inference memory for transformers up to 175B parameters by using mixed-precision decomposition: 8-bit matrix multiplication for 99.9
[Dettmers et al. 2023] Dettmers, Pagnoni, Holtzman, Zettlemoyer. QLoRA: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314. https://arxiv.org/abs/2305.14314
QLoRA enables finetuning of 65B-parameter LLMs on a single 48GB GPU by backpropagating through a frozen 4-bit quantized model into LoRA adapters, using NF4, Double Quantization, and Paged Optimizers.
[Dhariwal and Nichol 2021] Dhariwal, Nichol. Diffusion models beat gans on image synthesis. https://arxiv.org/abs/2105.05233
Dhariwal and Nichol show that diffusion models surpass GANs on image synthesis by improving UNet architecture and introducing classifier guidance to trade diversity for sample fidelity.
[Dillon et al. 2025] Dillon, Jaffe, Immorlica, Stanton. Shifting work patterns with generative AI. https://arxiv.org/abs/2504.11436
A six-month RCT of 6,000 workers finds AI access mainly changes individually controlled work patterns, reducing email time among users but not meeting time.
[Ding et al. 2026] Ding, Lyu, Kataria, Singh. Photonic rails in ML datacenters with opus. arXiv preprint arXiv:2602.12521. https://arxiv.org/abs/2602.12521
Opus replaces electrical packet switches in ML datacenter rail fabrics with optical circuit switches, using parallelism-driven reconfiguration to achieve over 23x network power reduction and 4x cost savings at under 6
[Dong et al. 2023] Dong, Xiong, Goyal, Zhang, Chow, Pan, Diao, Zhang, Shum, Zhang. RAFT: Reward rAnked FineTuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767. https://arxiv.org/abs/2304.06767
RAFT aligns generative models by iteratively sampling outputs, scoring them with a reward model, and fine-tuning only on the top-ranked subset, replacing PPO with a stable SFT-style loop.
[Dong et al. 2023] Dong, Lu, Zheng, Wu, Zhao, Tan, Huang, Hong, Wei, Chen. PUMA: Secure inference of LLaMA-7B in five minutes. arXiv preprint arXiv:2307.12533. https://arxiv.org/abs/2307.12533
The landmark secure multi-party computation result for LLMs: LLaMA-7B inference across three parties at roughly five minutes per generated token, defining the cost gap to trusted hardware.
[Dong et al. 2024] Dong, Ruan, Cai, Lai, Xu, Zhao, Chen. XGrammar: Flexible and efficient structured generation engine for large language models. arXiv preprint arXiv:2411.15100. https://arxiv.org/abs/2411.15100
XGrammar accelerates context-free grammar constrained decoding for LLMs by splitting vocabulary into context-independent and context-dependent tokens, achieving up to 100x per-token latency reduction.
[Dong et al. 2025] Dong, Mao, Ma, Bao, Chen, others. Agentic reinforced policy optimization. arXiv preprint arXiv:2507.19849. https://arxiv.org/abs/2507.19849
ARPO is an agentic RL algorithm for multi-turn LLM agents that uses entropy-based adaptive rollout to branch sampling at high-uncertainty tool-use steps, halving the tool-call budget needed by trajectory-level methods.
[Dosovitskiy et al. 2021] Dosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, Uszkoreit, Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. https://arxiv.org/abs/2010.11929
ViT shows that a pure Transformer applied directly to sequences of image patches matches or beats CNNs on image classification when pre-trained at sufficient scale.
[Dror et al. 2017] Dror, Baumer, Bogomolov, Reichart. Replicability analysis for natural language processing: Testing significance with multiple datasets. Transactions of the Association for Computational Linguistics. https://aclanthology.org/Q17-1033/
This paper proposes a replicability analysis framework for NLP comparisons across multiple datasets, addressing the multiple-comparison problem that arises in broad evaluation suites.
[Dror et al. 2018] Dror, Baumer, Shlomov, Reichart. The hitchhiker's guide to testing statistical significance in natural language processing. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. https://aclanthology.org/P18-1128/
Dror et al. survey statistical significance testing in NLP and recommend test choices that match common metrics and experimental designs.
[Du et al. 2022] Du, Huang, Dai, Tong, Lepikhin, Xu, Krikun, Zhou, Yu, Firat, Zoph, Fedus, Bosma, Zhou, Wang, Wang, Webster, Pellat, Robinson, Meier-Hellstern, Duke, Dixon, Zhang, Le, Wu, Chen, Cui. GLaM: Efficient scaling of language models with mixture-of-experts. https://arxiv.org/abs/2112.06905
GLaM scales a decoder-only language model to 1.2T parameters via sparsely activated MoE, matching or exceeding GPT-3 on 29 NLP tasks while using one-third the training energy.
[Du et al. 2025] Du, Tian, Ronanki, Rongali, Bodapati, Galstyan, Wells, Schwartz, Huerta, Peng. Context length alone hurts LLM performance despite perfect retrieval. https://arxiv.org/abs/2510.05381
Even with perfect retrieval, LLM reasoning performance degrades substantially (13.9
[Du 2026] Du. Tiered super-moore's law: Price evolution, production frontiers, and market competition in large language model inference services. https://arxiv.org/abs/2603.28576
Du documents rapid token-price decline across model tiers and a fall in inference-market concentration, framing token prices as a distinct digital-goods market.
[Dubois et al. 2024] Dubois, Galambosi, Liang, Hashimoto. Length-controlled AlpacaEval: a simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475. https://arxiv.org/abs/2404.04475
Length-Controlled AlpacaEval estimates judge preferences under a counterfactual zero length difference, reducing verbosity bias and improving correlation with Chatbot Arena rankings.
[Edge et al. 2024] Edge, Trinh, Cheng, Bradley, Chao, Mody, Truitt, Larson. From local to global: a graph RAG approach to query-focused summarization. arXiv preprint arXiv:2404.16130. https://arxiv.org/abs/2404.16130
GraphRAG builds an LLM-derived entity knowledge graph over a corpus, detects hierarchical communities, and uses map-reduce over community summaries to answer global sensemaking queries that vector RAG cannot handle.
[Efron 1979] Efron. Bootstrap methods: Another look at the jackknife. The Annals of Statistics. https://doi.org/10.1214/aos/1176344552
Efron's paper introduces the bootstrap as a general resampling method for estimating uncertainty without deriving a closed-form sampling distribution.
[Eigen AI Team and SGLang Team 2025] Eigen AI Team and SGLang Team. Accelerating SGLang with multiple token prediction. https://lmsys.org/blog/2025-07-17-mtp/
SGLang's multi-token-prediction speculative decoding for DeepSeek-V3 reports up to 60
[Eldan and Li 2023] Eldan, Li. TinyStories: How small can language models be and still speak coherent english?. https://arxiv.org/abs/2305.07759
TinyStories introduces a synthetic dataset of simple short stories to show that language models below 10 million parameters can generate fluent, coherent English text with emergent reasoning.
[Elhage et al. 2021] Elhage, Nanda, Olsson, Henighan, Joseph, Mann, Askell, Bai, Chen, Conerly, DasSarma, Drain, Ganguli, Hatfield-Dodds, Hernandez, Jones, Kernion, Lovitt, Ndousse, Amodei, Brown, Clark, Kaplan, McCandlish, Olah. A mathematical framework for transformer circuits. https://transformer-circuits.pub/2021/framework/index.html
This paper introduces a mathematical framework for mechanistic interpretability of transformers, reverse-engineering attention-only models via the residual stream, QK/OV circuits, and induction heads.
[Elhage et al. 2022] Elhage, Hume, Olsson, Schiefer, Henighan, Kravec, Hatfield-Dodds, Lasenby, Drain, Chen, Grosse, McCandlish, Kaplan, Amodei, Wattenberg, Olah. Toy models of superposition. https://arxiv.org/abs/2209.10652
Toy ReLU networks demonstrate that neural networks represent more features than dimensions via superposition, explaining polysemanticity and revealing a phase diagram governing when features align with neurons.
[Elo 1978] Elo. The rating of chessplayers, past and present. Arco Publishing. https://archive.org/details/ratingofchesspla0000eloa
Arpad Elo's 1978 book introduces the Elo rating system, a probabilistic method for ranking chess players by expected score from pairwise comparisons.
[Elsworth et al. 2025] Elsworth, Patterson, Dean, others. Measuring the environmental impact of delivering AI at google scale. arXiv preprint arXiv:2508.15734. https://arxiv.org/abs/2508.15734
Google measures the full-stack energy, carbon, and water footprint of Gemini Apps inference in production, finding the median text prompt consumes 0.24 Wh and showing a 44x emissions reduction over one year.
[Endsley 1995] Endsley. Toward a theory of situation awareness in dynamic systems. Human Factors. https://doi.org/10.1518/001872095779049543
Endsley formalizes situation awareness as a dynamic decision-making construct shaped by perception, comprehension, projection, workload, complexity, and automation.
[Enevoldsen and others 2025] Enevoldsen, others. MMTEB: Massive multilingual text embedding benchmark. https://arxiv.org/abs/2502.13595
MMTEB is a community-driven benchmark covering 500+ quality-controlled text embedding evaluation tasks across 250+ languages, with optimized subsets that reduce compute by 98
[Engels et al. 2025] Engels, Baek, Michaud, Tegmark. Scaling laws for scalable oversight. arXiv preprint arXiv:2504.18530. https://arxiv.org/abs/2504.18530
This paper proposes a quantitative framework for scalable oversight, modeling it as a game between capability-mismatched LLMs and deriving scaling laws and optimal Nested Scalable Oversight (NSO) configurations.
[Epoch AI 2026] Epoch AI. GDPval. https://epoch.ai/benchmarks/gdpval
Epoch AI's tracking of the GDPval benchmark, on which the best 2026 model wins or ties against human experts on roughly 71
[Erdil 2025] Erdil. Inference economics of language models. arXiv preprint arXiv:2506.04645. https://arxiv.org/abs/2506.04645
A theoretical model derives Pareto frontiers of serial token generation speed versus cost per token for LLM inference, accounting for arithmetic, HBM bandwidth, and network latency constraints across parallelism configurations.
[Eren et al. 2026] Eren, Krohn, Todorov. Financing the AI infrastructure boom: on- and off-balance sheet borrowing. https://www.bis.org/publ/qtrpdf/r_qt2603u.htm
A BIS Quarterly Review study documenting big-tech bond issuance passing $100B in 2025 and the special-purpose-vehicle shadow borrowing that creates new links between hyperscalers, private credit, and banks.
[Es et al. 2023] Es, James, Espinosa-Anke, Schockaert. RAGAS: Automated evaluation of retrieval augmented generation. arXiv preprint arXiv:2309.15217. https://arxiv.org/abs/2309.15217
RAGAS proposes reference-free metrics for RAG pipelines, including context relevance, context recall, faithfulness, and answer relevance.
[Esser et al. 2021] Esser, Rombach, Ommer. Taming transformers for high-resolution image synthesis. https://arxiv.org/abs/2012.09841
VQGAN combines a CNN-based discrete codebook with a transformer to enable autoregressive synthesis of high-resolution images at megapixel scale.
[Esser et al. 2024] Esser, Kulal, Blattmann, Entezari, Müller, Saini, Levi, Lorenz, Sauer, Boesel, Podell, Dockhorn, English, Lacey, Goodwin, Marek, Rombach. Scaling rectified flow transformers for high-resolution image synthesis. https://arxiv.org/abs/2403.03206
SD3 introduces improved noise sampling for rectified flow training and a new transformer architecture with bidirectional text-image token mixing, demonstrating predictable scaling for high-resolution text-to-image synthesis.
[Ethayarajh 2019] Ethayarajh. How contextual are contextualized word representations? Comparing the geometry of BERT, elmo, and GPT-2 embeddings. https://aclanthology.org/D19-1006/
This paper analyzes the geometry of contextualized word representations in BERT, ELMo, and GPT-2, finding they are anisotropic and that upper layers produce more context-specific embeddings.
[Ethayarajh et al. 2024] Ethayarajh, Xu, Muennighoff, Jurafsky, Kiela. KTO: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306. https://arxiv.org/abs/2402.01306
KTO aligns LLMs using a Kahneman-Tversky prospect theory objective that learns from binary desirability signals instead of preference pairs, matching or exceeding DPO at scales from 1B to 30B parameters.
[European Commission 2025] European Commission. The general-purpose AI code of practice. https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai
[European Commission 2025] European Commission. Guidelines for providers of general-purpose AI models. https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers
The Commission's guidance for GPAI providers under the AI Act: obligations apply from August 2, 2025, enforcement powers follow from August 2026, and providers must publish a training-content summary using the official template.
[European Commission 2025] European Commission. The general-purpose AI code of practice. https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai
The European Commission's GPAI Code of Practice gives providers a voluntary way to demonstrate AI Act compliance for transparency, copyright, and systemic-risk safety obligations.
[European Parliament and Council of the European Union 2024] European Parliament and Council of the European Union. Regulation (EU) 2024/1689 of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act). https://eur-lex.europa.eu/eli/reg/2024/1689/oj
[European Parliament and Council of the European Union 2024] European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
The EU AI Act requires high-risk AI systems to support human oversight, including awareness of automation bias, correct interpretation, override, and stop or interruption procedures.
[Ewen et al. 2025] Ewen, Dongen, Shilman. Durable AI loops: Fault tolerance across frameworks. https://www.restate.dev/blog/durable-ai-loops-fault-tolerance-across-frameworks-and-without-handcuffs
Restate's durable execution wraps agent loops to provide fault tolerance, resumability, observability, and human-in-the-loop support across any agent SDK without framework lock-in.
[Fast.io 2025] Fast.io. AI agent reliability report 2025. https://fast.io/resources/ai-agent-idempotent-operations/
A developer guide explaining why AI agents must use idempotent operations (atomic writes, idempotency keys, file locking) to prevent data corruption and duplicate side effects on retries.
[Faysse et al. 2024] Faysse, Sibille, Wu, Omrani, Viaud, Hudelot, Colombo. ColPali: Efficient document retrieval with vision language models. arXiv preprint arXiv:2407.01449. https://arxiv.org/abs/2407.01449
[Fedus et al. 2021] Fedus, Zoph, Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research. https://arxiv.org/abs/2101.03961
Switch Transformer simplifies MoE routing to a single expert per token, enabling trillion-parameter sparse models that achieve up to 7x pre-training speedup over T5 at equal FLOPs.
[Feng et al. 2023] Feng, Wan, Wen, McAleer, Wen, Zhang, Wang. Alphazero-like tree-search can guide large language model decoding and training. arXiv preprint arXiv:2309.17179. https://arxiv.org/abs/2309.17179
TS-LLM applies AlphaZero-like tree search with a learned value function to guide LLM decoding and iterative training across reasoning, planning, RLHF alignment, and decision-making tasks.
[Feng et al. 2025] Feng, Xue, Liu, An. Group-in-group policy optimization for LLM agent training. https://arxiv.org/abs/2505.10978
GiGPO proposes a two-level group-based RL algorithm for multi-turn LLM agent training that achieves fine-grained per-step credit assignment via anchor state grouping, without extra rollouts or critic models.
[Fernandes et al. 2023] Fernandes, Madaan, Liu, Farinhas, Martins, Bertsch, Souza, Zhou, Wu, Neubig, Martins. Bridging the gap: a survey on integrating human feedback for natural language generation. arXiv preprint arXiv:2305.00955. https://arxiv.org/abs/2305.00955
This survey formalizes human feedback for natural language generation, covering feedback formats, objectives, training use, decoding use, and the emerging role of AI feedback.
[Fielding et al. 2022] Fielding, Nottingham, Reschke. HTTP semantics (RFC 9110), section 15.5.3: 402 payment required. https://www.rfc-editor.org/rfc/rfc9110.html
The HTTP status code whose entire specification reads "reserved for future use," sitting in the protocol since HTTP/1.1 waiting for a payer without mental transaction costs.
[Financial Times 2025] Financial Times. Insurers seek to exclude AI liabilities from corporate policies. https://www.ft.com/content/ai-insurance-exclusions-2025
Reports that several large insurers filed to exclude AI liabilities from general corporate policies, the response of a market that cannot yet price correlated AI risk.
[FinOps Foundation 2026] FinOps Foundation. What is FinOps?. https://www.finops.org/introduction/what-is-finops/
FinOps defines technology cost management as a collaborative operating model across engineering, finance, product, and business teams.
[FinOps Open Cost and Usage Specification 2026] FinOps Open Cost and Usage Specification. FOCUS: FinOps open cost and usage specification. https://focus.finops.org/
FOCUS provides a vendor-neutral format for cost and usage data, making AI spend analyzable beside other technology spend; its virtual-currency columns cover token purchase and burn-down, with per-model token consumption scoped for the release after 1.4.
[Foundation for Defense of Democracies 2025] Foundation for Defense of Democracies. Rolling back export controls, U.S. offers china powerful AI chips. https://www.fdd.org/analysis/2025/12/10/rolling-back-export-controls-u-s-offers-china-powerful-ai-chips/
An FDD policy analysis arguing that the Trump administration's December 2025 rollback of AI chip export controls offers China access to powerful AI semiconductors, undermining prior restrictions.
[Frantar et al. 2022] Frantar, Ashkboos, Hoefler, Alistarh. GPTQ: Accurate post-training quantization for generative pre-trained transformers. https://arxiv.org/abs/2210.17323
GPTQ is a one-shot post-training quantization method that compresses GPT-scale models to 3-4 bits per weight in a few GPU hours with negligible accuracy loss, enabling OPT-175B inference on a single GPU.
[Friedl et al. 2026] Friedl, Ward, Rapoport, Everitt, Richens. The impossibility of eliciting latent knowledge. arXiv preprint arXiv:2606.12268. https://arxiv.org/abs/2606.12268
This paper formalizes eliciting latent knowledge with causal influence diagrams and proves that behavior-only feedback cannot guarantee honest agents in all cases.
[Friston 2010] Friston. The free-energy principle: a unified brain theory?. Nature Reviews Neuroscience.
Friston proposes the free-energy principle, arguing the brain minimizes a variational free-energy bound on surprise to perceive, learn, and act.
[Frostig et al. 2018] Frostig, Johnson, Leary. Compiling machine learning programs via high-level tracing. https://mlsys.org/Conferences/doc/2018/146.pdf
The four-page paper that defines JAX: trace pure Python/NumPy functions to a graph, lower to XLA, and expose differentiation, compilation, and vectorization as composable transformations.
[Fu et al. 2024] Fu, Bailis, Stoica, Zhang. Break the sequential dependency of LLM inference using lookahead decoding. https://arxiv.org/abs/2402.02057
Lookahead Decoding accelerates LLM autoregressive decoding without a draft model by generating and verifying n-grams in parallel via Jacobi iteration, achieving up to 4x speedup.
[Fu et al. 2025] Fu, Gao, Shen, Zhu, Mei, He, Xu, Wei, Mei, Wang, Yang, Yuan, Wu. AReaL: a large-scale asynchronous reinforcement learning system for language reasoning. arXiv preprint arXiv:2505.24298. https://arxiv.org/abs/2505.24298
AReaL is a fully asynchronous RL training system for large reasoning models that decouples generation from training, achieving up to 2.77x speedup over synchronous systems via a staleness-aware PPO variant.
[Furman 2025] Furman. Information-processing investment and US GDP growth (public remarks). https://x.com/jasonfurman/status/1971995367202775284
Furman's finding that information-processing equipment and software, about 4
[Gao et al. 2019] Gao, He, Tan, Qin, Wang, Liu. Representation degeneration problem in training natural language generation models. https://arxiv.org/abs/1907.12009
This paper identifies the representation degeneration problem, where word embeddings collapse into a narrow cone during likelihood-maximized training with weight tying, and proposes a regularization fix.
[Gao et al. 2020] Gao, Biderman, Black, Golding, Hoppe, Foster, Phang, He, Thite, Nabeshima, Presser, Leahy. The pile: An 800GB dataset of diverse text for language modeling. https://arxiv.org/abs/2101.00027
The Pile is an 825 GiB English text corpus assembled from 22 diverse sources, released to improve cross-domain generalization in large language model pretraining.
[Gao et al. 2021] Gao, Yao, Chen. SimCSE: Simple contrastive learning of sentence embeddings. https://arxiv.org/abs/2104.08821
SimCSE improves sentence embeddings via contrastive learning, using dropout noise as minimal data augmentation unsupervised and NLI entailment/contradiction pairs supervised.
[Gao et al. 2022] Gao, Schulman, Hilton. Scaling laws for reward model overoptimization. arXiv preprint arXiv:2210.10760. https://arxiv.org/abs/2210.10760
This paper measures reward model overoptimization in RLHF, deriving scaling laws showing how gold reward degrades as a function of KL divergence from the initial policy for both RL and best-of-n sampling.
[Gao et al. 2022] Gao, Madaan, Zhou, Alon, Liu, Yang, Callan, Neubig. PAL: Program-aided language models. arXiv preprint arXiv:2211.10435. https://arxiv.org/abs/2211.10435
Program-aided Language Models use LLMs to translate natural-language reasoning problems into executable programs, then offload computation to a Python interpreter.
[Gao et al. 2023] Gao, Yen, Yu, Chen. Enabling large language models to generate text with citations. arXiv preprint arXiv:2305.14627. https://arxiv.org/abs/2305.14627
ALCE evaluates answer correctness, fluency, and citation quality for retrieval-backed generation, showing that even strong models often lack complete citation support.
[Gao et al. 2024] Gao, Tour, Tillman, Goh, Troll, Radford, Sutskever, Leike, Wu. Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093. https://arxiv.org/abs/2406.04093
This paper proposes k-sparse autoencoders (TopK SAEs) that directly control sparsity to improve the reconstruction-sparsity frontier, establishes clean scaling laws for SAE size and sparsity, and trains a 16 million latent SAE on GPT-4 activations.
[Gao et al. 2025] Gao, Rajaram, Coxon, Govande, Baker, Mossing. Weight-sparse transformers have interpretable circuits. arXiv preprint arXiv:2511.13653. https://arxiv.org/abs/2511.13653
OpenAI researchers train transformers with mostly-zero weights so circuits are interpretable by construction, trading capability for interpretability at small scale.
[Gebru et al. 2021] Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III, Crawford. Datasheets for datasets. Communications of the ACM. https://arxiv.org/abs/1803.09010
Datasheets for Datasets proposes standardized dataset documentation covering motivation, composition, collection, preprocessing, uses, distribution, and maintenance.
[Gema et al. 2024] Gema, Leang, Hong, others. Are we done with MMLU?. arXiv preprint arXiv:2406.04127. https://arxiv.org/abs/2406.04127
A manual re-annotation of 5,700 MMLU questions estimates that 6.49
[Gemini Team, Google 2023] Gemini Team, Google. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. https://arxiv.org/abs/2312.11805
Google introduces Gemini, a multimodal model family (Ultra, Pro, Nano) jointly trained on text, image, audio, and video, with Gemini Ultra being the first model to exceed human-expert performance on MMLU.
[Gemini Team, Google 2024] Gemini Team, Google. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. https://arxiv.org/abs/2403.05530
Gemini 1.5 Pro and Flash are multimodal MoE models achieving near-perfect recall over up to 10M tokens of context across text, video, and audio.
[Gemma Team, Google DeepMind 2024] Gemma Team, Google DeepMind. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118. https://arxiv.org/abs/2408.00118
Gemma 2 introduces a family of open language models (2B, 9B, 27B) that use knowledge distillation, interleaved local-global attention, and GQA to match or outperform models 2-3x larger.
[Generalist AI 2025] Generalist AI. GEN-0: Embodied foundation models that scale with physical interaction. https://generalistai.com/blog/nov-04-2025-GEN-0
GEN-0 is a class of embodied foundation models for robotics that demonstrates scaling laws with real-world physical interaction data, trained on over 270,000 hours of manipulation data using a Harmonic Reasoning architecture for concurrent sensing and acting.
[Geng et al. 2025] Geng, Deng, Bai, Kolter, He. Mean flows for one-step generative modeling. arXiv preprint arXiv:2505.13447. https://arxiv.org/abs/2505.13447
[Geng and Neubig 2026] Geng, Neubig. Effective strategies for asynchronous software engineering agents. https://arxiv.org/abs/2603.21489
CAID introduces a multi-agent coordination paradigm using git worktrees, dependency graphs, and branch-and-merge integration to improve long-horizon software engineering agent performance by up to 26.7
[Gerstgrasser et al. 2024] Gerstgrasser, Schaeffer, Dey, Rafailov, Sleight, Hughes, Korbak, Agrawal, Pai, Gromov, Roberts, Yang, Donoho, Koyejo. Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. arXiv preprint arXiv:2404.01413. https://arxiv.org/abs/2404.01413
Accumulating real and synthetic data across model-fitting iterations provably avoids model collapse, while replacing old data with new synthetic data causes progressive performance degradation.
[ggml project 2023] ggml project. GGUF file format specification. https://github.com/ggml-org/ggml/blob/master/docs/gguf.md
The single-file format for local inference that packs tensors, tokenizer, and all metadata into one memory-mappable, quantization-aware file, needing no separate config.
[ggml-org 2023] ggml-org. llama.cpp. https://github.com/ggml-org/llama.cpp
llama.cpp is an open-source C/C++ library for running LLM inference locally, without GPU dependencies, using quantized weights.
[ggml-org 2023] ggml-org. GGUF. https://github.com/ggml-org/ggml/blob/master/docs/gguf.md
GGUF is a binary single-file format for storing LLM weights used by GGML-based inference engines, replacing GGML/GGJT with extensible key-value metadata and mmap-compatible layout.
[Ghazvininejad et al. 2019] Ghazvininejad, Levy, Liu, Zettlemoyer. Mask-predict: Parallel decoding of conditional masked language models. https://arxiv.org/abs/1904.09324
Mask-Predict introduces conditional masked language models with an iterative parallel decoding algorithm that achieves near-autoregressive translation quality while decoding significantly faster.
[Ghosal et al. 2025] Ghosal, Chakraborty, Reddy, Lu, Wang, Manocha, Huang, Ghavamzadeh, Bedi. Does thinking more always help? Understanding test-time scaling in reasoning models. arXiv preprint arXiv:2506.04210. https://arxiv.org/abs/2506.04210
This paper reports a non-monotonic test-time thinking curve: additional thinking first helps and then hurts, motivating parallel thinking and adaptive budget allocation.
[Ghosh et al. 2023] Ghosh, Hajishirzi, Schmidt. GenEval: An object-focused framework for evaluating text-to-image alignment. https://arxiv.org/abs/2310.11513
GENEVAL is an object-focused benchmark that uses object detection and discriminative vision models to evaluate compositional text-to-image alignment across tasks like counting, position, and attribute binding.
[Gitpod 2022] Gitpod. Data loss when workspace timeout fires before sync completes. https://github.com/gitpod-io/gitpod/issues/9544
A GitHub issue reporting that Gitpod workspace data is lost when a workspace times out and its pod enters a terminating state, because data is not synced to storage during the active session.
[Glassman et al. 1995] Glassman, Manasse, Abadi, Gauthier, Sobalvarro. The millicent protocol for inexpensive electronic commerce. https://www.w3.org/Conferences/WWW4/Papers/246/
DEC's broker-and-scrip design for transactions down to a tenth of a cent, a working answer to machine-granularity payment thirty years before there were machines to spend it.
[Glazer et al. 2024] Glazer, Erdil, Besiroglu, Chicharro, Chen, Gunning, Olsson, Denain, Ho, Oliveira Santos, Järviniemi, Barnett, Sandler, Vrzala, Sevilla, Ren, Pratt, Levine, Barkley, Stewart, Grechuk, Grechuk, Enugandla, Wildon. FrontierMath: a benchmark for evaluating advanced mathematical reasoning in AI. arXiv preprint arXiv:2411.04872. https://arxiv.org/abs/2411.04872
FrontierMath uses new, difficult, expert-vetted problems with automated verification to measure advanced mathematical reasoning while reducing contamination risk.
[Gloeckle et al. 2024] Gloeckle, Idrissi, Rozière, Lopez-Paz, Synnaeve. Better & faster large language models via multi-token prediction. https://arxiv.org/abs/2404.19737
Training LLMs with n independent output heads to predict multiple future tokens simultaneously improves sample efficiency, boosts coding benchmarks by  15
[Goldman Sachs 2026] Goldman Sachs. Tracking Trillions: the assumptions shaping the scale of the AI build-out. https://www.goldmansachs.com/insights/articles/tracking-trillions-the-assumptions-shaping-scale-of-the-ai-build-out
Places AI-related investment at roughly 1.5
[Gollwitzer 2025] Gollwitzer. Federal courts find fair use in AI training: Key takeaways from kadrey v. meta and bartz v. anthropic. https://www.jw.com/news/insights-kadrey-meta-bartz-anthropic-ai-copyright/
[Google 2025] Google. Private AI Compute: our next step in building private and helpful AI. https://blog.google/innovation-and-ai/products/google-private-ai-compute/
Google's confidential serving design runs Gemini models inside TPU enclaves with SEV-SNP on the CPU side, the second custom-silicon path to the same nobody-including-us claim.
[Google 2025] Google. Ironwood: The first google TPU for the age of inference. https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/
Google Ironwood is the seventh-generation TPU purpose-built for inference workloads, scaling to 9,216 chips with enhanced SparseCore, increased HBM capacity and bandwidth, and improved ICI networking.
[Google 2026] Google. Gemma 4: Byte for byte, the most capable open models. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/
[Google Cloud 2025] Google Cloud. Announcing the Agent2Agent protocol (A2A). https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/
Google introduced A2A as an open protocol for agents built by different vendors or frameworks to communicate, coordinate actions, and support long-running multi-agent workflows.
[Google DeepMind 2025] Google DeepMind. Gemini diffusion. https://deepmind.google/models/gemini-diffusion/
Gemini Diffusion is Google DeepMind's discrete diffusion language model that generates entire blocks of tokens at once, achieving 1479 tokens/sec with benchmark performance comparable to Gemini 2.0 Flash-Lite.
[Google DeepMind 2025] Google DeepMind. Veo. https://deepmind.google/models/veo/
[Google DeepMind 2025] Google DeepMind. Genie 3: a new frontier for world models. https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/
A real-time interactive world model generating 720p at 24fps with a short visual memory, turning a request from a batch job into a stateful per-user session with a tens-of-milliseconds frame deadline.
[Google DeepMind 2025] Google DeepMind. Gemini robotics 1.5: Pushing the frontier of generalist robots with embodied reasoning, thinking, and motion transfer. arXiv preprint arXiv:2510.03342. https://arxiv.org/abs/2510.03342
Gemini Robotics 1.5 introduces a multi-embodiment VLA with a Motion Transfer mechanism and embodied thinking that enables zero-shot skill transfer across different robot bodies.
[Google DeepMind 2025] Google DeepMind. Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/
[Google DeepMind 2026] Google DeepMind. Decoupled DiLoCo: Cross-datacenter training of language models. arXiv preprint arXiv:2604.21428. https://arxiv.org/abs/2604.21428
Decoupled DiLoCo extends the DiLoCo framework with fully asynchronous learners and a central synchronizer using minimum quorum and token-weighted merging to achieve zero-downtime pre-training under continuous hardware failures.
[Grattafiori and others 2024] Grattafiori, others. The llama 3 herd of models. https://arxiv.org/abs/2407.21783
Meta presents Llama 3, a herd of dense Transformer language models at 8B, 70B, and 405B parameters trained on 15T tokens, achieving quality comparable to GPT-4 across diverse tasks.
[Graves et al. 2006] Graves, Fernández, Gomez, Schmidhuber. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. https://www.cs.toronto.edu/~graves/icml_2006.pdf
CTC introduces a training method for RNNs that labels unsegmented input sequences directly, eliminating the need for pre-segmentation or post-processing.
[Graves 2012] Graves. Sequence transduction with recurrent neural networks. ICML 2012 Representation Learning Workshop; arXiv:1211.3711. https://arxiv.org/abs/1211.3711
Graves 2012 introduces the RNN Transducer (RNN-T), an end-to-end probabilistic sequence transduction system that jointly models input-output and output-output dependencies without requiring a pre-defined alignment.
[Gray 2019] Gray. Getting started with CUDA graphs. https://developer.nvidia.com/blog/cuda-graphs/
[Greenblatt et al. 2024] Greenblatt, Shlegeris, Sachan, Roger. AI control: Improving safety despite intentional subversion. arXiv preprint arXiv:2312.06942. https://arxiv.org/abs/2312.06942
AI control evaluates whether safety protocols can remain acceptable when a powerful model may intentionally subvert the task.
[Greenblatt et al. 2024] Greenblatt, Denison, Wright, Roger, MacDiarmid, Marks, Treutlein, Belonax, Chen, Duvenaud, Khan, Michael, Mindermann, Perez, Petrini, Uesato, Kaplan, Shlegeris, Bowman, Hubinger. Alignment faking in large language models. arXiv preprint arXiv:2412.14093. https://arxiv.org/abs/2412.14093
Greenblatt et al. show Claude 3 Opus, given only situational information about being retrained, selectively complies during training to preserve its preferences, with alignment-faking reasoning rising from 14
[Greshake et al. 2023] Greshake, Abdelnabi, Mishra, Endres, Holz, Fritz. Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. https://arxiv.org/abs/2302.12173
This paper introduces indirect prompt injection, an attack where adversaries plant malicious instructions in data retrieved by LLM-integrated applications to remotely hijack model behavior without direct access.
[Greshake et al. 2023] Greshake, Abdelnabi, Mishra, Endres, Holz, Fritz. Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. arXiv preprint arXiv:2302.12173. https://arxiv.org/abs/2302.12173
This paper introduces indirect prompt injection, an attack where adversaries plant malicious instructions in data retrieved by LLM-integrated applications to remotely hijack the LLM without direct user access.
[Griewank 1992] Griewank. Achieving logarithmic growth of temporal and spatial complexity in reverse automatic differentiation. Optimization Methods and Software. https://doi.org/10.1080/10556789208805505
[Griewank and Walther 2008] Griewank, Walther. Evaluating derivatives: Principles and techniques of algorithmic differentiation. SIAM. https://epubs.siam.org/doi/book/10.1137/1.9780898717761
The canonical reference on algorithmic differentiation: the cheap gradient principle, the complexity rules, and provably optimal checkpointing schedules.
[Griewank 2012] Griewank. Who invented the reverse mode of differentiation?. Documenta Mathematica, Extra Volume ISMP. https://ftp.gwdg.de/pub/EMIS/journals/DMJDMV/vol-ismp/52_griewank-andreas-b.pdf
A readable history of reverse-mode differentiation from Linnainmaa's 1970 insight to modern practice, including the cheap gradient principle and checkpointing complexity results.
[Grinsztajn et al. 2022] Grinsztajn, Oyallon, Varoquaux. Why do tree-based models still outperform deep learning on typical tabular data?. https://arxiv.org/abs/2207.08815
A careful benchmark showing tree ensembles remain state of the art on medium-size tabular data, and tracing the advantage to inductive bias: robustness to uninformative features and a fit for non-smooth, axis-aligned target functions.
[Grünwald 2004] Grünwald. A tutorial introduction to the minimum description length principle. arXiv preprint math/0406077. https://arxiv.org/abs/math/0406077
Grünwald's tutorial introduces Rissanen's MDL principle, framing statistical model selection and inductive inference as data compression to automatically guard against overfitting.
[Gu et al. 2018] Gu, Bradbury, Xiong, Li, Socher. Non-autoregressive neural machine translation. https://arxiv.org/abs/1711.02281
The Non-Autoregressive Transformer (NAT) generates all output tokens in parallel using fertility predictions as a latent variable, achieving order-of-magnitude lower inference latency at a cost of about 2.0 BLEU points versus an autoregressive teacher.
[Gu et al. 2019] Gu, Wang, Zhao. Levenshtein transformer. https://arxiv.org/abs/1905.11006
Levenshtein Transformer (LevT) is a partially autoregressive sequence generation model that uses insertion and deletion operations with dual-policy imitation learning, achieving comparable translation quality with up to 5x decoding speedup.
[Gu et al. 2021] Gu, Goel, Ré. Efficiently modeling long sequences with structured state spaces. https://arxiv.org/abs/2111.00396
S4 introduces a structured parameterization of SSMs via low-rank plus normal decomposition, enabling O(N+L) computation and state-of-the-art performance on long-range sequence benchmarks including the previously unsolved Path-X task.
[Gu and Dao 2023] Gu, Dao. Mamba: Linear-time sequence modeling with selective state spaces. https://arxiv.org/abs/2312.00752
Mamba introduces selective SSMs with input-dependent parameters and a hardware-aware parallel scan, achieving Transformer-quality language modeling with linear-time inference and training.
[Guan et al. 2024] Guan, Joglekar, Wallace, Jain, Barak, Helyar, Dias, Vallone, Ren, Wei, Chung, Toyer, Heidecke, Beutel, Glaese. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339. https://arxiv.org/abs/2412.16339
Deliberative Alignment teaches models safety specifications and trains them to reason over those specifications before answering, improving jailbreak robustness while reducing over-refusal.
[Guha et al. 2025] Guha, Marten, Keh, others. OpenThoughts: Data recipes for reasoning models. arXiv preprint arXiv:2506.04178. https://arxiv.org/abs/2506.04178
Over a thousand controlled ablations of the reasoning-data pipeline produce OpenThoughts3-1.2M; OpenThinker3-7B trained on it reaches 53
[Gulati et al. 2020] Gulati, Qin, Chiu, Parmar, Zhang, Yu, Han, Wang, Zhang, Wu, Pang. Conformer: Convolution-augmented transformer for speech recognition. https://arxiv.org/abs/2005.08100
Conformer combines convolution and self-attention in a single encoder block to capture both local and global audio features, achieving state-of-the-art ASR results on LibriSpeech.
[Gulrajani and Hashimoto 2023] Gulrajani, Hashimoto. Likelihood-based diffusion language models. https://arxiv.org/abs/2305.18619
Plaid 1B is a 1-billion-parameter likelihood-based diffusion language model that outperforms GPT-2 124M on zero-shot perplexity benchmarks via algorithmic improvements and compute-optimal scaling laws.
[Gunasekar et al. 2023] Gunasekar, Zhang, Aneja, Mendes, Del Giorno, Gopi, Javaheripi, Kauffmann, Rosa, Saarikivi, Salim, Shah, Behl, Wang, Bubeck, Eldan, Kalai, Lee, Li. Textbooks are all you need. https://arxiv.org/abs/2306.11644
phi-1, a 1.3B-parameter code LLM trained on only 7B tokens of textbook-quality data, achieves 50.6
[Gunjal et al. 2025] Gunjal, Wang, Lau, Nath, others. Rubrics as rewards: Reinforcement learning beyond verifiable domains. arXiv preprint arXiv:2507.17746. https://arxiv.org/abs/2507.17746
Rubrics as Rewards decomposes open-ended judgments into per-criterion, checklist-style rubrics graded by a model and used as the reward for on-policy RL, outperforming LLM-as-judge Likert baselines on HealthBench and GPQA-Diamond.
[Guo et al. 2017] Guo, Pleiss, Sun, Weinberger. On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning. https://proceedings.mlr.press/v70/guo17a.html
Guo et al. show that modern neural networks can be accurate yet miscalibrated, and that simple temperature scaling can substantially improve probabilistic calibration.
[Guo et al. 2024] Guo, Xia, Yu, Ao, Huang. LightRAG: Simple and fast retrieval-augmented generation. arXiv preprint arXiv:2410.05779. https://arxiv.org/abs/2410.05779
LightRAG integrates graph-based text indexing with a dual-level retrieval system into RAG, enabling more coherent responses to complex queries by capturing entity relationships and supporting incremental knowledge-base updates.
[Guo et al. 2025] Guo, Yang, Zhang, Song, Zhang, Xu, Zhu, Ma, Wang, Bi, others. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature. https://www.nature.com/articles/s41586-025-09422-z
DeepSeek-R1 shows that large-scale RL with verifiable rewards can elicit long reasoning behavior, with R1-Zero using RL without supervised cold start and R1 adding multi-stage training for readability and stability.
[Gupta et al. 2020] Gupta, Wu, Wang, Naumov, Reagen, Brooks, Cottel, Hazelwood, Hempstead, Jia, Lee, Malevich, Mudigere, Smelyanskiy, Xiong, Zhang. The architectural implications of facebook's DNN-based personalized recommendation. https://arxiv.org/abs/1906.03109
[Gururangan et al. 2020] Gururangan, Marasović, Swayamdipta, Lo, Beltagy, Downey, Smith. Don't stop pretraining: Adapt language models to domains and tasks. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. https://arxiv.org/abs/2004.10964
Shows that additional pretraining on domain and task corpora improves downstream performance across biomedical, computer science, news, and review tasks.
[Haas et al. 2025] Haas, Yona, D'Antonio, Goldshtein, Das. SimpleQA verified: a reliable factuality benchmark to measure parametric knowledge. arXiv preprint arXiv:2509.07968. https://arxiv.org/abs/2509.07968
SimpleQA Verified is a 1,000-prompt revision of SimpleQA that de-duplicates, rebalances topics, and reconciles sources to remove incorrect labels from the grading set.
[Haldar et al. 2019] Haldar, Abdool, Ramanathan, Xu, Yang, Duan, Zhang, Barrow-Williams, Turnbull, Collins, Legrand. Applying deep learning to airbnb search. https://arxiv.org/abs/1810.09591
[Hales et al. 2015] Hales, Adams, Bauer, Dang, Harrison, Hoang, Kaliszyk, Magron, McLaughlin, Nguyen, Nguyen, Nipkow, Obua, Pleso, Rute, Solovyev, Ta, Tran, Trieu, Urban, Vu, Zumkeller. A formal proof of the kepler conjecture. arXiv preprint arXiv:1501.02155. https://arxiv.org/abs/1501.02155
The Flyspeck project gives a formal proof of the Kepler conjecture using HOL Light and Isabelle, showing how a major mathematical result can be checked by proof assistants.
[Han et al. 2024] Han, Rao, Ettinger, Jiang, Lin, Lambert, Choi, Dziri. WildGuard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. https://arxiv.org/abs/2406.18495
WILDGUARD is an open lightweight LLM safety moderation tool that jointly detects prompt harmfulness, response harmfulness, and model refusal across 13 risk categories, including adversarial jailbreaks.
[Hashimoto and Liang 2025] Hashimoto, Liang. CS336: Language modeling from scratch. https://cs336.stanford.edu/spring2025/
Stanford's implementation-heavy language-modeling course walks through tokenizer construction, Transformer implementation, systems optimization, scaling laws, data processing, evaluation, and alignment.
[Hayes and Krippendorff 2007] Hayes, Krippendorff. Answering the call for a standard reliability measure for coding data. Communication Methods and Measures. https://doi.org/10.1080/19312450709336664
Hayes and Krippendorff argue for Krippendorff's alpha as a general reliability coefficient that handles many coders, missing data, and different measurement levels.
[He et al. 2019] He, Sainath, Prabhavalkar, McGraw, others. Streaming end-to-end speech recognition for mobile devices. https://arxiv.org/abs/1811.06621
Google presents a streaming end-to-end RNN-T speech recognizer for mobile devices that outperforms a conventional CTC model in both latency and WER on voice search and dictation tasks.
[He 2019] He. The state of machine learning frameworks in 2019. https://thegradient.pub/state-of-ml-frameworks-2019-pytorch-dominates-research-tensorflow-dominates-industry/
Settles the framework war with counted conference papers instead of vibes: PyTorch at 69
[He 2022] He. Making deep learning go brrrr from first principles. https://horace.io/brrr_intro.html
The gentlest correct introduction to GPU performance: every workload is memory-bound, compute-bound, or overhead-bound, and fusion is the most important optimization a deep-learning compiler performs.
[He 2025] He. Defeating nondeterminism in LLM inference. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
He traces LLM inference nondeterminism to reduction kernels whose numerics depend on dynamic batch size rather than to concurrency alone, and demonstrates batch-invariant kernels that make temperature-zero decoding bit-identical at roughly 1.6 to 2x decode time.
[Hebb 1949] Hebb. The organization of behavior: a neuropsychological theory. John Wiley & Sons.
Hebb proposes that learning strengthens synapses between co-active neurons (cells that fire together wire together), the basis of Hebbian learning.
[Heer 2019] Heer. Agency plus automation: Designing artificial intelligence into interactive systems. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.1807184115
Heer argues for interactive systems that use predictive models to augment human work through shared representations, review, revision, and dismissal.
[Hendrycks and Gimpel 2016] Hendrycks, Gimpel. Gaussian error linear units (gelus). https://arxiv.org/abs/1606.08415
This paper introduces GELU, an activation function defined as x times the Gaussian CDF, which outperforms ReLU and ELU across vision, NLP, and speech tasks.
[Hendrycks et al. 2020] Hendrycks, Burns, Basart, Zou, Mazeika, Song, Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300. https://arxiv.org/abs/2009.03300
MMLU introduces a 57-subject multiple-choice benchmark spanning STEM, humanities, and social sciences to measure text models' breadth of world knowledge in zero-shot and few-shot settings.
[Hendrycks et al. 2021] Hendrycks, Burns, Kadavath, Arora, Basart, Tang, Song, Steinhardt. Measuring mathematical problem solving with the MATH dataset. arXiv preprint arXiv:2103.03874. https://arxiv.org/abs/2103.03874
The MATH dataset introduces 12,500 competition mathematics problems with step-by-step solutions, showing that large Transformer scaling alone cannot solve challenging mathematical reasoning.
[Hermann and Del Balso 2017] Hermann, Del Balso. Meet michelangelo: Uber's machine learning platform. https://www.uber.com/blog/michelangelo-machine-learning-platform/
[Hernández-Orallo and others 2026] Hernández-Orallo, others. General scales unlock AI evaluation with explanatory and predictive power. arXiv preprint arXiv:2503.06378. https://arxiv.org/abs/2503.06378
This paper proposes 18 general demand-level rubrics (ADeLe) to annotate LLM benchmark instances, enabling both explanatory ability profiles and instance-level performance prediction in- and out-of-distribution.
[Hines et al. 2024] Hines, Lopez, Hall, Zarfati, Zunger, Kiciman. Defending against indirect prompt injection attacks with spotlighting. arXiv preprint arXiv:2403.14720. https://arxiv.org/abs/2403.14720
Spotlighting is a prompt engineering defense against indirect prompt injection attacks that transforms untrusted input text to signal its provenance, reducing attack success rate from over 50
[Ho et al. 2020] Ho, Jain, Abbeel. Denoising diffusion probabilistic models. https://arxiv.org/abs/2006.11239
Ho et al. show that DDPM trained with a denoising score matching objective achieves high-quality image synthesis, reaching FID 3.17 on CIFAR10.
[Ho and Salimans 2022] Ho, Salimans. Classifier-free diffusion guidance. https://arxiv.org/abs/2207.12598
Classifier-free guidance (CFG) jointly trains a conditional and unconditional diffusion model, mixing their score estimates at inference to trade off sample quality and diversity without requiring a separate classifier.
[Ho et al. 2022] Ho, Salimans, Gritsenko, Chan, Norouzi, Fleet. Video diffusion models. https://arxiv.org/abs/2204.03458
Ho et al. extend image diffusion models to video via a space-time factorized 3D U-Net and reconstruction-guided conditional sampling, achieving state-of-the-art results on video generation and prediction benchmarks.
[Hoffmann et al. 2022] Hoffmann, Borgeaud, Mensch, Buchatskaya, Cai, Rutherford, Las Casas, Hendricks, Welbl, Clark, Hennigan, Noland, Millican, Driessche, Damoc, Guy, Osindero, Simonyan, Elsen, Rae, Vinyals, Sifre. Training compute-optimal large language models. https://arxiv.org/abs/2203.15556
Shows model size and training tokens should scale equally (about 20 tokens per parameter), so most large models are badly undertrained; the compute-matched 70B Chinchilla beats much larger models like Gopher and GPT-3.
[Hofstätter et al. 2020] Hofstätter, Althammer, Schröder, Sertkan, Hanbury. Improving efficient neural ranking models with cross-architecture knowledge distillation. arXiv preprint arXiv:2010.02666. https://arxiv.org/abs/2010.02666
Margin-MSE is a cross-architecture knowledge distillation loss that trains efficient passage ranking models (TK, ColBERT, PreTT, BERT_DOT) to match the score margins of a BERT_CAT teacher, improving retrieval effectiveness without sacrificing query latency.
[Hofstätter et al. 2021] Hofstätter, Lin, Yang, Lin, Hanbury. Efficiently teaching an effective dense retriever with balanced topic aware sampling. https://arxiv.org/abs/2104.06967
TAS-Balanced trains a BERT_DOT dual-encoder dense retrieval model on a single consumer GPU by composing batches from topic clusters and balancing pairwise margin scores with dual-teacher knowledge distillation.
[Hohnhold et al. 2015] Hohnhold, O'Brien, Tang. Focusing on the long-term: It's good for users and business. https://research.google/pubs/focusing-on-the-long-term-its-good-for-users-and-business/
[Hollmann et al. 2025] Hollmann, Müller, Purucker, Krishnakumar, Körfer, Hoo, Schirrmeister, Hutter. Accurate predictions on small data with a tabular foundation model. Nature. https://www.nature.com/articles/s41586-024-08328-6
TabPFN, a transformer pretrained on millions of synthetic tabular tasks, reads a whole training table as context and beats tuned tree ensembles on datasets up to about ten thousand samples, in seconds.
[Hong et al. 2024] Hong, Lee, Thorne. ORPO: Monolithic preference optimization without reference model. arXiv preprint arXiv:2403.07691. https://arxiv.org/abs/2403.07691
ORPO is a monolithic preference alignment algorithm that merges SFT and preference optimization into one step using an odds ratio penalty, eliminating the need for a reference model.
[Hong et al. 2025] Hong, Troynikov, Huber. Context rot: How increasing input tokens impacts LLM performance. https://www.trychroma.com/research/context-rot
Chroma's "Context Rot" report evaluates 18 LLMs and shows that performance on controlled retrieval tasks degrades non-uniformly as input 词元 count grows, even when task complexity is held constant.
[Hooker 2021] Hooker. The hardware lottery. Communications of the ACM. https://arxiv.org/abs/2009.06489
A research idea can win because it suits the available software and hardware rather than because it is superior; as silicon specializes around matmuls, ideas off that path pay escalating costs even to be evaluated.
[Hooker 2026] Hooker. On the slow death of scaling. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5877662
[Horvitz 1999] Horvitz. Principles of mixed-initiative user interfaces. https://doi.org/10.1145/302979.303030
Horvitz frames user interfaces as mixed-initiative systems in which humans and computers negotiate control rather than forcing either direct manipulation or full automation.
[Houlsby et al. 2019] Houlsby, Giurgiu, Jastrzebski, Morrone, Laroussilhe, Gesmundo, Attariyan, Gelly. Parameter-efficient transfer learning for NLP. https://arxiv.org/abs/1902.00751
Adapter modules inserted into BERT layers match full fine-tuning performance on GLUE while training only 3.6
[Hsieh et al. 2024] Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, Ginsburg. RULER: What's the real context size of your long-context language models?. arXiv preprint arXiv:2404.06654. https://arxiv.org/abs/2404.06654
RULER is a synthetic benchmark with flexible context length and task complexity that reveals most long-context LLMs fail well below their claimed context sizes on tasks beyond simple retrieval.
[Hsu et al. 2021] Hsu, Bolte, Tsai, Lakhotia, Salakhutdinov, Mohamed. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing. https://arxiv.org/abs/2106.07447
HuBERT learns speech representations by predicting offline k-means cluster assignments on masked regions, matching or surpassing wav2vec 2.0 on ASR benchmarks.
[Hu et al. 2021] Hu, Shen, Wallis, Allen-Zhu, Li, Wang, Wang, Chen. LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685. https://arxiv.org/abs/2106.09685
LoRA freezes pretrained weights and injects trainable low-rank matrix pairs into each Transformer layer, reducing trainable parameters by 10,000x and GPU memory by 3x versus full fine-tuning with no added inference latency.
[Hu et al. 2022] Hu, Shen, Wallis, Allen-Zhu, Li, Wang, Wang, Chen. LoRA: Low-rank adaptation of large language models. https://arxiv.org/abs/2106.09685
LoRA freezes pretrained weights and injects trainable low-rank decomposition matrices into Transformer layers, reducing trainable parameters by 10,000x with no added inference latency.
[Hu et al. 2024] Hu, Tu, Han, He, Cui, Long, Zheng, Fang, Huang, Zhao, Zhang, Thai, Zhang, Wang, Yao, Zhao, Zhou, Cai, Zhai, Ding, Jia, Zeng, Li, Liu, Sun. MiniCPM: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395. https://arxiv.org/abs/2404.06395
Presents small (1.2B/2.4B) models that rival 7B-13B LLMs, using model wind-tunnel scaling experiments and a Warmup-Stable-Decay learning-rate schedule that enables continuous training.
[Hu et al. 2024] Hu, Wang, Fang, Fu, Cheng, Yu. ELLA: Equip diffusion models with LLM for enhanced semantic alignment. arXiv preprint arXiv:2403.05135. https://arxiv.org/abs/2403.05135
ELLA is a lightweight LLM adapter for CLIP-based text-to-image diffusion models that uses a Timestep-Aware Semantic Connector to improve dense prompt following without retraining the U-Net or LLM.
[Hu and others 2024] Hu, others. OpenRLHF: An easy-to-use, scalable and high-performance RLHF framework. arXiv preprint arXiv:2405.11143. https://arxiv.org/abs/2405.11143
OpenRLHF is an open-source RLHF and RLVR training framework built on Ray, vLLM, and DeepSpeed that achieves 1.22x to 1.68x speedups over verl while requiring fewer lines of code.
[Hu et al. 2025] Hu, Luschi, Balanca. Elucidating the design space of FP4 training. arXiv preprint arXiv:2509.17791. https://arxiv.org/abs/2509.17791
Maps the design space of 4-bit training across block-scaled formats, Hadamard transforms, and stochastic rounding, identifying which combinations keep fp4 matmuls near baseline quality at acceptable overhead.
[Huang et al. 2019] Huang, Cheng, Bapna, Firat, Chen, Chen, Lee, Ngiam, Le, Wu, Chen. GPipe: Efficient training of giant neural networks using pipeline parallelism. https://arxiv.org/abs/1811.06965
GPipe proposes pipeline parallelism via micro-batch splitting to scale neural networks beyond single-accelerator memory limits with near-linear speedup across multiple accelerators.
[Huang et al. 2020] Huang, Sharma, Sun, Xia, Zhang, Pronin, Padmanabhan, Ottaviano, Yang. Embedding-based retrieval in facebook search. https://arxiv.org/abs/2006.11632
The full production recipe for embedding retrieval: two-tower model, ANN index, hybrid dense-plus-keyword retrieval, and hard-negative mining, published four years before the RAG boom rediscovered each item.
[Huang et al. 2024] Huang, Zhang, Shan, He. Compression represents intelligence linearly. https://arxiv.org/abs/2404.09937
Across 31 public LLMs and 12 benchmarks, average benchmark scores correlate almost linearly with the models' compression efficiency on external text corpora, with Pearson coefficients around -0.95 per ability area.
[Huang et al. 2024] Huang, Chen, Mishra, Zheng, Yu, Song, Zhou. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798. https://arxiv.org/abs/2310.01798
LLMs cannot self-correct reasoning without external feedback; intrinsic self-correction consistently degrades accuracy on reasoning benchmarks across GPT-3.5, GPT-4, and Llama-2.
[Huang et al. 2025] Huang, Yu, Wang, others. R-zero: Self-evolving reasoning LLM from zero data. arXiv preprint arXiv:2508.05004. https://arxiv.org/abs/2508.05004
R-Zero initializes a Challenger and a Solver from one base model and co-evolves them, the Challenger proposing ever harder tasks and the Solver learning to solve them, generating its own curriculum from zero external data.
[Huawei 2025] Huawei. Groundbreaking SuperPoD interconnect: Leading a new paradigm for AI infrastructure. https://www.huawei.com/en/news/2025/9/hc-xu-keynote-speech
Huawei's announcement of the Atlas 950 SuperPoD, an optically interconnected scale-up domain of 8,192 Ascend chips slated for Q4 2026, with self-developed HBM.
[Hubert et al. 2025] Hubert, Mehta, Sartran, others. Olympiad-level formal mathematical reasoning with reinforcement learning. Nature. https://www.nature.com/articles/s41586-025-09833-y
AlphaProof trains an AlphaZero-style RL agent to write Lean proofs, using millions of auto-formalized problems, and reached silver-medal standard on the 2024 IMO problems; autoformalization is the limiting step.
[Hubinger et al. 2024] Hubinger, Denison, Mu, Lambert, Tong, MacDiarmid, Lanham, Ziegler, Maxwell, Cheng, Jermyn, Askell, Radhakrishnan, Anil, Duvenaud, Ganguli, Barez, Clark, Ndousse, others. Sleeper agents: Training deceptive llms that persist through safety training. arXiv preprint arXiv:2401.05566. https://arxiv.org/abs/2401.05566
Backdoors implanted in LLMs via deliberate training persist through supervised fine-tuning, RL, and adversarial training, suggesting standard safety techniques cannot reliably remove deceptive alignment.
[Hugging Face 2023] Hugging Face. Audit shows that safetensors is safe and ready to become the default. https://huggingface.co/blog/safetensors-security-audit
The safetensors format and its commissioned external audit, which found no arbitrary-code-execution path: a load is a bounds-checked parse and a memory-map, with no reconstruction hook.
[Hugging Face 2026] Hugging Face. The state of open source on hugging face: Spring 2026. https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026
The hub's own report of roughly two million public models in early 2026, with a vanishing fraction of models accounting for about half of all downloads.
[Hugging Face Nanotron Team 2025] Hugging Face Nanotron Team. The ultra-scale playbook: Training llms on GPU clusters. https://huggingface.co/spaces/nanotron/ultrascale-playbook
A Hugging Face Nanotron guide focused specifically on training large language models on large GPU clusters, with an accompanying PDF and interactive web version.
[Hui et al. 2024] Hui, Yang, Cui, Yang, Liu, Zhang, Liu, Zhang, Yu, Lu, Dang, Fan, Zhang, Yang, Men, Huang, Zheng, Miao, Quan, Feng, Ren, Ren, Zhou, Lin. Qwen2.5-coder technical report. arXiv preprint arXiv:2409.12186. https://arxiv.org/abs/2409.12186
Reports Qwen2.5-Coder, a code-specialized family further trained on more than 5.5T tokens with cleaned code, scalable synthetic data, and balanced mixing while retaining general and math skills.
[Hunton Andrews Kurth 2025] Hunton Andrews Kurth. Bartz v. anthropic: Settlement reached after landmark summary judgment and class certification. https://www.insidetechlaw.com/blog/2025/09/bartz-v-anthropic-settlement-reached-after-landmark-summary-judgment-and-class-certification
[Huyen 2025] Huyen. AI engineering: Building applications with foundation models. O'Reilly Media. https://www.oreilly.com/library/view/ai-engineering/9781098166298/
Chip Huyen's book covers AI engineering: building production applications on top of foundation models, including evaluation, adaptation techniques, and serving.
[Hwang et al. 2025] Hwang, Wang, Gu. Dynamic chunking for end-to-end hierarchical sequence modeling. arXiv preprint arXiv:2507.07955. https://arxiv.org/abs/2507.07955
H-Net replaces the tokenization pipeline with a hierarchical model that learns content-dependent chunk boundaries end to end, outperforming byte-level and BPE Transformers at matched compute.
[IETF 2025] IETF. OAuth 2.0 on-behalf-of user for AI agents. https://datatracker.ietf.org/doc/draft-oauth-ai-agents-on-behalf-of-user/
An IETF Internet-Draft defining an OAuth 2.0 extension with a new grant type and requested_agent parameter so AI agents can obtain delegated access tokens on behalf of users with explicit consent.
[Ilharco et al. 2023] Ilharco, Ribeiro, Wortsman, Gururangan, Schmidt, Hajishirzi, Farhadi. Editing models with task arithmetic. https://arxiv.org/abs/2212.04089
Task arithmetic introduces task vectors, directions in weight space obtained by subtracting pre-trained from fine-tuned weights, which can be negated or added to edit model behavior without retraining.
[Inan et al. 2023] Inan, Upasani, Chi, Rungta, Iyer, Mao, Tontchev, Hu, Fuller, Testuggine, Khabsa. Llama guard: LLM-based input-output safeguard for human-AI conversations. arXiv preprint arXiv:2312.06674. https://arxiv.org/abs/2312.06674
Llama Guard is an instruction-tuned Llama2-7b model that classifies both user prompts and LLM responses against a customizable safety risk taxonomy, matching or exceeding existing content moderation APIs.
[Inception Labs et al. 2025] Inception Labs, Khanna, Kharbanda, Li, others. Mercury: Ultra-fast language models based on diffusion. arXiv preprint arXiv:2506.17298. https://arxiv.org/abs/2506.17298
Mercury Coder is a family of commercial-scale diffusion LLMs that generate tokens in parallel to achieve up to 10x higher throughput than autoregressive models while maintaining comparable code quality.
[International Energy Agency 2025] International Energy Agency. Energy and AI. https://www.iea.org/reports/energy-and-ai
[International Energy Agency 2025] International Energy Agency. Energy and AI. https://www.iea.org/reports/energy-and-ai
The IEA report analyzes how AI increases data-centre electricity demand while also affecting power systems, energy security, emissions, innovation, and affordability.
[International Organization for Standardization 2025] International Organization for Standardization. ISO/DIS 22144: Authenticity of information, content credentials. https://www.iso.org/standard/90726.html
[Invariant Labs 2025] Invariant Labs. GitHub MCP exploited: Accessing private repositories via MCP. https://invariantlabs.ai/blog/mcp-github-vulnerability
Invariant Labs demonstrated a critical prompt injection vulnerability in the official GitHub MCP server that lets attackers leak private repository data via malicious GitHub Issues.
[Irving et al. 2018] Irving, Christiano, Amodei. AI safety via debate. arXiv preprint arXiv:1805.00899. https://arxiv.org/abs/1805.00899
Debate trains agents to argue opposing sides so a human judge can use the adversarial exchange to evaluate claims too hard to inspect directly.
[ISO/IEC 2023] ISO/IEC. ISO/IEC 42001:2023, information technology, artificial intelligence, management system. https://www.iso.org/standard/42001
The first international AI management-system standard: it certifies that an organization has documented policies, risk assessments, roles, and an improvement loop, not that any model behaves safely.
[Ivanov et al. 2021] Ivanov, Dryden, Ben-Nun, Li, Hoefler. Data movement is all you need: a case study on optimizing transformers. Proceedings of Machine Learning and Systems (MLSys). https://arxiv.org/abs/2007.00072
[Izacard et al. 2022] Izacard, Caron, Hosseini, Riedel, Bojanowski, Joulin, Grave. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research (TMLR). https://arxiv.org/abs/2112.09118
Contriever trains an unsupervised dense retriever with contrastive learning on unlabeled text, matching or outperforming BM25 on 11 of 15 BEIR benchmark datasets without any labeled data.
[Jacovi and others 2025] Jacovi, others. The FACTS grounding leaderboard: Benchmarking llms' ability to ground responses to long-form input. arXiv preprint arXiv:2501.03200. https://arxiv.org/abs/2501.03200
FACTS Grounding scores long-form responses for full grounding in a supplied document of up to 32k tokens, using a panel of judge models from different families that separately check eligibility and factual grounding.
[Jagadeesh et al. 2026] Jagadeesh, Arora, Saab, Malik, Trofimov, Tsimpourlas, Heidecke, Singhal. Reinforcement learning towards broadly and persistently beneficial models. https://alignment.openai.com/beneficial-rl/
OpenAI reports that reinforcement learning on realistic scenarios targeting seven beneficial traits (truthfulness, epistemic humility, metacognitive transparency, corrigibility, risk sensitivity, universal fairness, and concern for human welfare) improved alignment on most of 53 held-out benchmarks, transferred from health-only training to other domains, and persisted under adversarial steering and fine-tuning.
[Jamshidi et al. 2025] Jamshidi, Nafi, Dakhel, Shahabi, Khomh, Ezzati-Jivan. Securing the model context protocol: Defending llms against tool poisoning and adversarial attacks. https://arxiv.org/abs/2512.06556
This paper studies MCP semantic attacks such as tool poisoning, shadowing, and rug pulls, and proposes descriptor signing, semantic vetting, and runtime guardrails.
[JFrog Security Research 2024] JFrog Security Research. Data scientists targeted by malicious hugging face ML models with silent backdoor. https://jfrog.com/blog/data-scientists-targeted-by-malicious-hugging-face-ml-models-with-silent-backdoor/
A scan finding roughly a hundred genuinely malicious models on the largest model hub, including one whose pickle reconstruction hook opened a reverse shell to an attacker on load.
[Jiang et al. 2024] Jiang, Sablayrolles, Roux, Mensch, Savary, Bamford, Chaplot, Casas, Hanna, Bressand, Lengyel, Bour, Lample, Lavaud, Saulnier, Lachaux, Stock, Subramanian, Yang, Antoniak, Scao, Gervet, Lavril, Wang, Lacroix, Sayed. Mixtral of experts. https://arxiv.org/abs/2401.04088
Mixtral 8x7B is a sparse MoE decoder-only model with 46.7B total parameters that activates only 12.9B per token, outperforming Llama 2 70B with 6x faster inference under an Apache 2.0 license.
[Jimenez et al. 2023] Jimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can language models resolve real-world GitHub issues?. arXiv preprint arXiv:2310.06770. https://arxiv.org/abs/2310.06770
SWE-bench is a benchmark of 2,294 real GitHub issue-resolution tasks across 12 Python repositories, where models must edit codebases to pass tests, and top models like Claude 2 solve only 1.96
[Jin et al. 2025] Jin, Zeng, Yue, Yoon, Arik, Wang, Zamani, Han. Search-R1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516. https://arxiv.org/abs/2503.09516
SEARCH-R1 trains LLMs via reinforcement learning to autonomously issue multi-turn search queries during reasoning, outperforming RAG baselines by up to 24
[Johnson et al. 2019] Johnson, Douze, Jégou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data. https://arxiv.org/abs/1702.08734
[Jones et al. 2020] Jones, Nadalin, Campbell, Bradley, Mortimore. OAuth 2.0 token exchange. https://datatracker.ietf.org/doc/html/rfc8693
RFC 8693 specifies the OAuth 2.0 Token Exchange protocol, defining how clients request security tokens from an authorization server to support impersonation and delegation flows.
[Jordan et al. 2024] Jordan, Jin, Boza, You, Cesista, Newhouse, Bernstein. Muon: An optimizer for hidden layers in neural networks. https://kellerjordan.github.io/posts/muon/
Introduces Muon, an optimizer for hidden layers that orthogonalizes momentum-based updates, setting training-speed records on NanoGPT and CIFAR-10.
[Jouppi et al. 2017] Jouppi, Young, Patil, Patterson, others. In-datacenter performance analysis of a tensor processing unit. https://arxiv.org/abs/1704.04760
[Ju et al. 2024] Ju, Wang, Shen, Tan, Xin, Yang, Liu, Leng, Song, Tang, others. NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models. https://arxiv.org/abs/2403.03100
NaturalSpeech 3 proposes a factorized codec (FACodec) and factorized diffusion models that decompose speech into disentangled subspaces for zero-shot TTS, achieving human-level naturalness on multi-speaker data.
[Kahneman 2011] Kahneman. Thinking, fast and slow. Farrar, Straus,Giroux.
Kahneman contrasts fast intuitive System 1 thinking with slow deliberate System 2 reasoning, and the cognitive biases each produces.
[Kalai et al. 2025] Kalai, Nachum, Vempala, Zhang. Why language models hallucinate. arXiv preprint arXiv:2509.04664. https://arxiv.org/abs/2509.04664
This paper argues that LLM hallucinations arise from statistical pressures during pretraining (via a reduction to binary classification) and persist because most benchmarks reward guessing over expressing uncertainty.
[Kang et al. 2025] Kang, Yin, Gao, others. How far is video generation from world model: a physical law perspective. https://openreview.net/forum?id=DLlVjZQ7vD
Scaling diffusion-based video generation models achieves perfect in-distribution generalization but fails at out-of-distribution physical law extrapolation, showing scaling alone cannot uncover fundamental physical laws.
[Kang et al. 2025] Kang, Chen, Han, Inan, Wutschitz, Chen, Sim, Rajmohan. ACON: Optimizing context compression for long-horizon LLM agents. https://arxiv.org/abs/2510.00615
ACON is a model-agnostic framework that optimizes LLM agent context compression via natural language guideline refinement, reducing peak token usage by 26–54
[Kantamneni et al. 2025] Kantamneni, Engels, Rajamanoharan, Tegmark, Nanda. Are sparse autoencoders useful? A case study in sparse probing. arXiv preprint arXiv:2502.16681. https://arxiv.org/abs/2502.16681
A probing study across 113 datasets finds that sparse autoencoder (SAE) latents fail to consistently outperform simple baselines for LLM activation probing under data scarcity, class imbalance, label noise, or covariate shift.
[Kaplan et al. 2020] Kaplan, McCandlish, Henighan, Brown, Chess, Child, Gray, Radford, Wu, Amodei. Scaling laws for neural language models. https://arxiv.org/abs/2001.08361
Establishes that language-model loss falls as a power law in model size, dataset size, and compute, and that compute-optimal training favors very large models trained on relatively little data, stopped before convergence.
[Kaplan 2023] Kaplan. Hardware VM isolation in the cloud. ACM Queue. https://dl.acm.org/doi/abs/10.1145/3623392
The architect of AMD SEV recounts a decade of iterating VM isolation designs against real attacks, a candid record of how confidential computing actually hardened.
[Kapoor et al. 2024] Kapoor, Bommasani, Klyman, Longpre, Ramaswami, Cihon, Hopkins, Bankston, Biderman, Bogen, Chowdhury, Engler, Henderson, Jernite, Lazar, Maffulli, Nelson, Pineau, Skowron, Song, Storchan, Zhang, Ho, Liang, Narayanan. On the societal impact of open foundation models. https://arxiv.org/abs/2403.07918
This position paper analyzes open-weight foundation models through their benefits, risks, and marginal risk relative to existing technologies, clarifying where evidence is still missing.
[Karpathy 2025] Karpathy. +1 for “context engineering” over “prompt engineering”. https://x.com/karpathy/status/1937902205765607626
[Karpinska et al. 2021] Karpinska, Akoury, Iyyer. The perils of using mechanical turk to evaluate open-ended text generation. arXiv preprint arXiv:2109.06835. https://arxiv.org/abs/2109.06835
Karpinska et al. show that crowd workers can fail to distinguish human and model-written open-ended stories, while expert teachers and paired examples provide stronger evaluation signals.
[Karpukhin et al. 2020] Karpukhin, Oğuz, Min, Lewis, Wu, Edunov, Chen, Yih. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906. https://arxiv.org/abs/2004.04906
DPR shows that a dual-encoder fine-tuned on question-passage pairs using in-batch negatives outperforms BM25 by 9-19
[Karras et al. 2022] Karras, Aittala, Aila, Laine. Elucidating the design space of diffusion-based generative models. https://arxiv.org/abs/2206.00364
Karras et al. decompose diffusion model training and sampling into a clean design space, identifying improvements to preconditioning and schedules that set new FID records on CIFAR-10 and ImageNet-64.
[Kawrakow 2023] Kawrakow. k-quants. https://github.com/ggml-org/llama.cpp/pull/1684
This llama.cpp pull request introduces k-quants, a series of 2-6 bit quantization methods with quantization mixes, implemented for scalar, AVX2, ARM_NEON, and CUDA backends.
[Khalifa et al. 2025] Khalifa, Agarwal, Logeswaran, others. Process reward models that think. arXiv preprint arXiv:2504.16828. https://arxiv.org/abs/2504.16828
ThinkPRM is a generative PRM that verbalizes step-by-step verification as a chain of thought, outperforming discriminative PRMs and LLM-as-judge while training on roughly 1
[Khan et al. 2024] Khan, Hughes, Valentine, Ruis, Sachan, Radhakrishnan, Grefenstette, Bowman, Rocktäschel, Perez. Debating with more persuasive llms leads to more truthful answers. arXiv preprint arXiv:2402.06782. https://arxiv.org/abs/2402.06782
Empirical study showing that LLM debate, where two expert models argue opposing answers for a weaker judge, enables scalable oversight by achieving 76
[Khashabi et al. 2021] Khashabi, Stanovsky, Bragg, Lourie, Kasai, Choi, Smith, Weld. GENIE: Toward reproducible and standardized human evaluation for text generation. arXiv preprint arXiv:2101.06561. https://arxiv.org/abs/2101.06561
GENIE studies human-evaluation design choices for text generation and introduces a standardized platform that improves reproducibility across tasks and annotator populations.
[Khattab and Zaharia 2020] Khattab, Zaharia. ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. arXiv preprint arXiv:2004.12832. https://arxiv.org/abs/2004.12832
ColBERT introduces a late interaction architecture that independently encodes queries and documents with BERT, then scores relevance via MaxSim, achieving competitive effectiveness at two orders of magnitude lower latency.
[Kim and Rush 2016] Kim, Rush. Sequence-level knowledge distillation. https://arxiv.org/abs/1606.07947
Sequence-level knowledge distillation trains a small student NMT model on teacher beam-search outputs, yielding a 10x faster student that matches teacher BLEU with greedy decoding.
[Kim et al. 2024] Kim, Pertsch, others. OpenVLA: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246. https://arxiv.org/abs/2406.09246
OpenVLA is a 7B-parameter open-source VLA trained on 970k robot demonstrations that outperforms the closed 55B RT-2-X model and supports LoRA fine-tuning on consumer GPUs.
[Kim et al. 2025] Kim, Kotha, Liang, Hashimoto. Pre-training under infinite compute. arXiv preprint arXiv:2509.14786. https://arxiv.org/abs/2509.14786
Under data-constrained, compute-unlimited pre-training, combining aggressive regularization, parameter scaling, and ensemble scaling achieves a lower loss asymptote and 5.17x data efficiency over the standard recipe.
[Kimi Team 2025] Kimi Team. Kimi K2: Open agentic intelligence. arXiv preprint arXiv:2507.20534. https://arxiv.org/abs/2507.20534
Kimi K2 is a 1T-total, 32B-active MoE model pre-trained on 15.5T tokens with MuonClip, a Muon variant whose QK-clip removes attention-logit blowups, with zero loss spikes over the run.
[Kimi Team et al. 2025] Kimi Team, Du, Gao, Xing, Jiang, Chen, Li, Xiao, Du, Liao, Tang, Wang, Zhang, others. Kimi k1.5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599. https://arxiv.org/abs/2501.12599
Kimi k1.5 reports a multimodal RL reasoning system that scales long-context RL without MCTS, value functions, or PRMs, and introduces long2short methods for short-CoT models.
[Kirchenbauer et al. 2023] Kirchenbauer, Geiping, Wen, Katz, Miers, Goldstein. A watermark for large language models. https://arxiv.org/abs/2301.10226
This paper proposes a statistical watermarking framework for LLM output that embeds detectable signals into generated text by biasing token sampling toward a randomized "green list," detectable from as few as 25 tokens without model access.
[Kirk et al. 2024] Kirk, Vidgen, Röttger, Hale. The benefits, risks and bounds of personalizing the alignment of large language models to individuals. Nature Machine Intelligence. https://www.nature.com/articles/s42256-024-00820-y
The normative treatment of personalized alignment: what tuning a model to an individual buys, the profiling and bias-reinforcement risks it carries, and the case for explicit bounds.
[Kleppmann 2016] Kleppmann. How to do distributed locking. https://martin.kleppmann.com/2016/02/08/how-to-do-distributed-locking.html
Martin Kleppmann argues that Redis's Redlock algorithm is unsafe for correctness-critical distributed locking due to timing assumptions that process pauses and clock drift can violate.
[Kohavi et al. 2020] Kohavi, Tang, Xu. Trustworthy online controlled experiments: a practical guide to a/B testing. Cambridge University Press. https://experimentguide.com/
The reference on running online controlled experiments at scale, distilled from tens of thousands of experiments a year at Microsoft, Google, and LinkedIn; read the chapters on overall evaluation criteria and long-term metrics first.
[Kojima et al. 2022] Kojima, Gu, Reid, Matsuo, Iwasawa. Large language models are zero-shot reasoners. https://arxiv.org/abs/2205.11916
Appending "Let's think step by step" to any question (Zero-shot-CoT) substantially improves LLM reasoning across arithmetic and symbolic tasks without any task-specific few-shot examples.
[Komatsuzaki et al. 2022] Komatsuzaki, Puigcerver, Lee-Thorp, Ruiz, Mustafa, Ainslie, Tay, Dehghani, Houlsby. Sparse upcycling: Training mixture-of-experts from dense checkpoints. https://arxiv.org/abs/2212.05055
Sparse upcycling initializes a MoE model from a pretrained dense checkpoint, outperforming both dense continuation and MoE training from scratch at roughly 50
[Koomey et al. 2025] Koomey, Schmidt, Das. Electricity demand growth and data centers: a guide for the perplexed. https://bipartisanpolicy.org/report/electricity-demand-growth-and-data-centers/
[Köpf et al. 2023] Köpf, Kilcher, Rütte, Anagnostidis, Tam, Stevens, Barhoum, Duc, Stanley, Nagyfi, Shahul, Suri, Glushkov, Dantuluri, Maguire, Schuhmann, Nguyen, Mattick. OpenAssistant conversations: Democratizing large language model alignment. arXiv preprint arXiv:2304.07327. https://arxiv.org/abs/2304.07327
OpenAssistant Conversations releases a permissively licensed, crowd-sourced alignment corpus with 161,443 messages in 35 languages, 461,292 ratings, and over 10,000 annotated conversation trees.
[Korthikanti et al. 2022] Korthikanti, Casper, Lym, McAfee, Andersch, Shoeybi, Catanzaro. Reducing activation recomputation in large transformer models. https://arxiv.org/abs/2205.05198
This paper introduces sequence parallelism and selective activation recomputation to reduce activation memory by 5x and cut activation recomputation overhead by over 90
[Kubernetes 2023] Kubernetes. PVC binding prevents pod rescheduling. https://github.com/kubernetes/kubernetes/issues/121436
A Kubernetes bug report where a pod that fails at the bind stage leaves its PVC and PV permanently bound to a node, causing the pod to remain stuck in Pending forever.
[Kubernetes 2026] Kubernetes. Multi-tenancy. https://kubernetes.io/docs/concepts/security/multi-tenancy/
Kubernetes documents the isolation and fairness tools that AI platforms inherit and extend for model, cache, index, and tool boundaries.
[Kudo 2018] Kudo. Subword regularization: Improving neural network translation models with multiple subword candidates. https://aclanthology.org/P18-1007/
Subword regularization trains NMT models on multiple probabilistically sampled subword segmentations and introduces a unigram language model segmentation algorithm as a probabilistic alternative to BPE.
[Kudo and Richardson 2018] Kudo, Richardson. SentencePiece: a simple and language independent subword tokenizer and detokenizer for neural text processing. https://aclanthology.org/D18-2012/
SentencePiece is a language-independent subword tokenizer and detokenizer that trains directly from raw sentences using BPE or unigram language model, enabling purely end-to-end text processing without pre-tokenization.
[Kumar et al. 2023] Kumar, Seetharaman, Luebs, Kumar, Kumar. High-fidelity audio compression with improved RVQGAN. https://arxiv.org/abs/2306.06546
Kumar et al. present Improved RVQGAN, a universal neural audio codec that compresses 44.1 kHz audio to discrete tokens at 8 kbps (90x compression) with higher fidelity than EnCodec and SoundStream.
[Kumar et al. 2025] Kumar, Roh, Naseh, Karpinska, Iyyer, Houmansadr, Bagdasarian. OverThink: Slowdown attacks on reasoning llms. arXiv preprint arXiv:2502.02542. https://arxiv.org/abs/2502.02542
OverThink shows that decoy reasoning problems in retrieved content can make reasoning LLMs spend many more reasoning tokens while still returning contextually correct answers.
[Kusupati and others 2022] Kusupati, others. Matryoshka representation learning. https://arxiv.org/abs/2205.13147
Matryoshka Representation Learning (MRL) trains a single embedding vector so any prefix of its dimensions forms an accurate representation, enabling up to 14x smaller embeddings at equal accuracy for classification and retrieval.
[Kutasov et al. 2025] Kutasov, Sun, Colognese, Weij, Petrini, Zhang, Hughes, Deng, Sleight, Tracy, Shlegeris, Benton. SHADE-arena: Evaluating sabotage and monitoring in LLM agents. arXiv preprint arXiv:2506.15740. https://arxiv.org/abs/2506.15740
SHADE-Arena pairs benign long-horizon agent tasks with hidden sabotage side-tasks and scores both the agent's covert success and an LLM monitor's ability to flag the transcript.
[Kwa et al. 2025] Kwa, West, Becker, others. Measuring AI ability to complete long software tasks. arXiv preprint arXiv:2503.14499. https://arxiv.org/abs/2503.14499
METR scores an agent by the human task duration at which it succeeds 50
[Kwa et al. 2025] Kwa, West, others. Measuring AI ability to complete long tasks. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
METR proposes measuring AI progress by the length of tasks agents can autonomously complete, finding this horizon has doubled roughly every 7 months over the past 6 years.
[Kwon et al. 2023] Kwon, Li, Zhuang, Sheng, Zheng, Yu, Gonzalez, Zhang, Stoica. Efficient memory management for large language model serving with PagedAttention. https://arxiv.org/abs/2309.06180
PagedAttention manages LLM KV cache in non-contiguous paged blocks, reducing fragmentation and enabling sharing, which vLLM uses to improve serving throughput 2-4x.
[Kwon et al. 2023] Kwon, Li, Zhuang, Sheng, Zheng, Yu, Gonzalez, Zhang, Stoica. Efficient memory management for large language model serving with PagedAttention. https://arxiv.org/abs/2309.06180
PagedAttention manages KV cache in fixed-size non-contiguous pages inspired by OS virtual memory, enabling vLLM to serve LLMs at 2-4x higher throughput than prior systems.
[Lambert et al. 2024] Lambert, Morrison, Pyatkin, Huang, Ivison, Brahman, Miranda, Liu, Dziri, Lyu, Gu, Malik, Graf, Hwang, Yang, Bras, Tafjord, Wilhelm, Soldaini, Smith, Wang, Dasigi, Hajishirzi. Tulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124. https://arxiv.org/abs/2411.15124
Tulu 3 is a fully open post-training recipe for Llama 3.1 base models, combining SFT, DPO, and RLVR with released data, weights, and training code.
[LangChain 2025] LangChain. LangGraph time travel and branching. https://docs.langchain.com/oss/python/langgraph/use-time-travel
LangGraph's time-travel feature lets users replay past graph executions and fork from any prior checkpoint to explore alternative execution paths.
[LangChain 2025] LangChain. LangGraph interrupts. https://docs.langchain.com/oss/python/langgraph/interrupts
LangGraph's interrupt() API lets agent graph nodes pause mid-execution and wait for human input, resuming via Command(resume=...) on the same thread.
[Lattner et al. 2021] Lattner, Amini, Bondhugula, Cohen, Davis, Pienaar, Riddle, Shpeisman, Vasilache, Zinenko. MLIR: Scaling compiler infrastructure for domain specific computation. https://arxiv.org/abs/2002.11054
MLIR is the compiler framework of composable dialects and progressive lowering that XLA, Triton's internals, Mosaic, and most newer ML compilers now share.
[Le et al. 2023] Le, Vyas, Shi, Karrer, Sari, Moritz, Williamson, Manohar, Adi, Mahadeokar, Hsu. Voicebox: Text-guided multilingual universal speech generation at scale. https://arxiv.org/abs/2306.15687
Voicebox is a non-autoregressive flow-matching model trained on 50K+ hours of speech to perform zero-shot TTS, denoising, and content editing via in-context learning, outperforming VALL-E on both intelligibility and audio similarity.
[Lechner 2025] Lechner. Why we started with JAX but moved to PyTorch. https://mlechner.substack.com/p/why-we-started-with-jax-but-moved
[Lee and See 2004] Lee, See. Trust in automation: Designing for appropriate reliance. Human Factors. https://doi.org/10.1518/hfes.46.1.50_30392
Lee and See review trust in automation as a basis for appropriate reliance under complexity, with display design and context shaping whether reliance is justified.
[Lee et al. 2022] Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini. Deduplicating training data makes language models better. https://aclanthology.org/2022.acl-long.577/
Deduplicating language model training data reduces verbatim memorization tenfold and lowers train-test overlap, yielding better accuracy with fewer training steps.
[Lee et al. 2025] Lee, Roy, Xu, Raiman, Shoeybi, Catanzaro, Ping. NV-embed: Improved techniques for training llms as generalist embedding models. https://arxiv.org/abs/2405.17428
NV-Embed improves decoder-only LLMs as generalist embedding models via a latent attention pooling layer, bidirectional contrastive training, and two-stage instruction tuning, reaching top MTEB scores.
[Lee et al. 2025] Lee, Chen, Dua, Cer, others. Gemini embedding: Generalizable embeddings from gemini. arXiv preprint arXiv:2503.07891. https://arxiv.org/abs/2503.07891
Gemini Embedding is a multilingual text embedding model initialized from Gemini that achieves state-of-the-art results on MMTEB across 250+ languages, English, and code benchmarks.
[Lees et al. 2022] Lees, Tran, Tay, Sorensen, Gupta, Metzler, Vasserman. A new generation of perspective API: Efficient multilingual character-level transformers. https://arxiv.org/abs/2202.11176
This paper presents UTC, a single vocabulary-free Charformer-based multilingual Transformer for toxic content classification that powers the next generation of Google Jigsaw's Perspective API.
[Lepikhin et al. 2020] Lepikhin, Lee, Xu, Chen, Firat, Huang, Krikun, Shazeer, Chen. GShard: Scaling giant models with conditional computation and automatic sharding. https://arxiv.org/abs/2006.16668
GShard introduces lightweight annotation APIs and an XLA compiler extension enabling automatic SPMD sharding of a 600B-parameter MoE Transformer trained on 2048 TPU v3 devices for multilingual translation across 100 languages.
[Leroy 2009] Leroy. A formally verified compiler back-end. Journal of Automated Reasoning. https://arxiv.org/abs/0902.2137
CompCert's verified back end proves semantic preservation from Cminor to PowerPC assembly in Coq, making compiler correctness part of the trusted evidence chain.
[Letta 2025] Letta. Letta: Stateful agents with memory as a first-class primitive. https://www.letta.com/
Letta is an AI research lab and product platform for building stateful agents that maintain persistent memory, learn continuously, and improve over time, originating from the MemGPT project.
[Letta 2025] Letta. Benchmarking AI agent memory. https://www.letta.com/blog/benchmarking-ai-agent-memory
A Letta agent using only filesystem tools scores 74.0
[Leviathan et al. 2023] Leviathan, Kalman, Matias. Fast inference from transformers via speculative decoding. https://arxiv.org/abs/2211.17192
Speculative decoding uses a small draft model to propose tokens verified in parallel by the large target model, achieving 2X-3X speedup without changing the output distribution.
[Leviathan et al. 2023] Leviathan, Kalman, Matias. Fast inference from transformers via speculative decoding. https://arxiv.org/abs/2211.17192
Speculative decoding uses a small draft model to propose tokens in parallel, then verifies them with the target model, achieving 2-3x inference speedup with identical output distribution and no retraining.
[Levine 2026] Levine. Agentic commerce won't kill cards, but it'll open a gap. https://a16zcrypto.com/posts/article/agentic-commerce-wont-kill-cards/
Argues agents will not optimize away card rails, which bundle credit, pre-authorization, and fraud guarantees, while stablecoins take the merchants and machine services no processor will underwrite.
[Lewis et al. 2020] Lewis, Perez, Piktus, Petroni, Karpukhin, Goyal, Küttler, Lewis, Yih, Rocktäschel, Riedel, Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. arXiv preprint arXiv:2005.11401. https://arxiv.org/abs/2005.11401
RAG combines a pre-trained parametric seq2seq model with a dense non-parametric Wikipedia index retrieved via DPR, achieving state-of-the-art on open-domain QA and producing more factual generation than parametric-only baselines.
[Lewis et al. 2021] Lewis, Bhosale, Dettmers, Goyal, Zettlemoyer. BASE layers: Simplifying training of large, sparse models. https://arxiv.org/abs/2103.16716
BASE layers replace MoE routing heuristics and auxiliary balancing losses by formulating token-to-expert assignment as a linear assignment problem, guaranteeing equal load across experts with no new hyperparameters.
[Li et al. 2020] Li, Zhou, He, Wang, Yang, Li. On the sentence embeddings from pre-trained language models. https://arxiv.org/abs/2011.05864
BERT-flow applies normalizing flows to remap BERT's anisotropic sentence embedding space to an isotropic Gaussian, improving semantic textual similarity without supervised fine-tuning.
[Li et al. 2022] Li, Thickstun, Gulrajani, Liang, Hashimoto. Diffusion-LM improves controllable text generation. https://arxiv.org/abs/2205.14217
Diffusion-LM adapts continuous diffusion models to text by iteratively denoising Gaussian vectors into word vectors, enabling gradient-based plug-and-play controllable generation over complex attributes like syntactic structure.
[Li et al. 2022] Li, Li, Dall, Gu, Nieh, Sait, Stockwell. Design and verification of the arm confidential compute architecture. https://www.usenix.org/conference/osdi22/presentation/li
Arm's realm-based confidential architecture, designed together with a formal verification of its firmware, the field's cleanest paper design still waiting for volume hardware.
[Li et al. 2023] Li, Li, Savarese, Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. https://arxiv.org/abs/2301.12597
BLIP-2 introduces a lightweight Querying Transformer (Q-Former) that bridges frozen image encoders and frozen LLMs in two pre-training stages, achieving strong vision-language performance with far fewer trainable parameters.
[Li et al. 2023] Li, Zhang, Zhang, Long, Xie, Zhang. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281. https://arxiv.org/abs/2308.03281
GTE is a 110M-parameter general text embedding model trained with multi-stage contrastive learning on open-source data that outperforms OpenAI's embedding API on the MTEB benchmark.
[Li et al. 2023] Li, Zhao, Yu, Song, Li, Yu, Li, Huang, Li. API-bank: a comprehensive benchmark for tool-augmented llms. arXiv preprint arXiv:2304.08244. https://arxiv.org/abs/2304.08244
API-Bank evaluates tool-augmented language models with runnable API tools and dialogues that test planning, API retrieval, and API calling.
[Li et al. 2024] Li, Wei, Zhang, Zhang. EAGLE: Speculative sampling requires rethinking feature uncertainty. https://arxiv.org/abs/2401.15077
EAGLE is a speculative decoding framework that predicts second-to-top-layer features conditioned on a one-step-ahead token sequence, achieving 2.7x-3.5x latency speedup over autoregressive LLM decoding with lossless output distribution.
[Li et al. 2024] Li, Wei, Zhang, Zhang. EAGLE-2: Faster inference of language models with dynamic draft trees. https://arxiv.org/abs/2406.16858
EAGLE-2 improves speculative decoding by replacing static draft trees with context-aware dynamic draft trees, achieving 3.05x-4.26x lossless LLM inference speedup without additional training.
[Li et al. 2024] Li, Huang, Yang, Venkitesh, Locatelli, Ye, Cai, Lewis, Chen. SnapKV: LLM knows what you are looking for before generation. https://arxiv.org/abs/2404.14469
SnapKV is a fine-tuning-free method that compresses the KV cache for long-context LLMs by identifying important attention positions from an observation window before generation, achieving 3.6x faster decoding and 8.2x memory reduction at 16K tokens.
[Li et al. 2024] Li, Cai, Cao, Zhang, Cai, Bai, Jia, Li, Han. DistriFusion: Distributed parallel inference for high-resolution diffusion models. https://arxiv.org/abs/2402.19481
Splits a single high-resolution diffusion sample across GPUs via displaced patch parallelism, reusing the previous step's feature maps so workers communicate asynchronously, up to 6.1x lower latency with no quality loss.
[Li et al. 2024] Li, Angelopoulos, Chiang. Does style matter? Disentangling style and substance in chatbot arena. https://lmsys.org/blog/2024-08-28-style-control/
[Li et al. 2025] Li, Wei, Zhang, Zhang. EAGLE-3: Scaling up inference acceleration of large language models via training-time test. https://arxiv.org/abs/2503.01840
EAGLE-3 replaces feature prediction with direct token prediction and multi-layer feature fusion via training-time test, enabling a scaling law for speculative decoding that achieves up to 6.5x inference speedup.
[Li et al. 2025] Li, Meng, Lin, Luo, Tian, Ma, Huang, Chua. ScreenSpot-pro: GUI grounding for professional high-resolution computer use. arXiv preprint arXiv:2504.07981. https://arxiv.org/abs/2504.07981
Grounding on 4K professional software: the best model of its moment located 18.9
[Liang et al. 2022] Liang, Zhang, Kwon, Yeung, Zou. Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning. https://arxiv.org/abs/2203.02053
This paper identifies and explains "modality gap," a geometric phenomenon where image and text embeddings in multimodal models like CLIP occupy separate regions of the shared representation space due to the cone effect and contrastive learning.
[Liang et al. 2022] Liang, Bommasani, Lee, Tsipras, Soylu, Yasunaga, Zhang, Narayanan, Wu, Kumar, Newman, Yuan, Yan, Zhang, Cosgrove, Manning, Ré, Acosta-Navas, Hudson, Zelikman, Durmus, Ladhak, Rong, Ren, Yao, Wang, Santhanam, Orr, Zheng, Yuksekgonul, Suzgun, Kim, Guha, Chatterji, Khattab, Henderson, Huang, Chi, Xie, Santurkar, Ganguli, Hashimoto, Icard, Zhang, Chaudhary, Wang, Li, Mai, Zhang, Koreeda. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110. https://arxiv.org/abs/2211.09110
HELM benchmarks 30 language models across 42 scenarios with 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency) to expose trade-offs invisible to single-metric evaluation.
[Liang et al. 2022] Liang, Bommasani, Lee, Tsipras, Soylu, Yasunaga, Zhang, Narayanan, Wu, Kumar, Newman, Yuan, Yan, Zhang, Cosgrove, Manning, Ré, Acosta-Navas, Hudson, others. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110. https://arxiv.org/abs/2211.09110
HELM benchmarks 30 language models across 42 scenarios and 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency) under standardized conditions to improve transparency in LLM evaluation.
[Liang et al. 2025] Liang, Niu, Li, Zhang, Wang, Xiong, Fan, Tang, Song, Wang, Yang. SafeRAG: Benchmarking security in retrieval-augmented generation of large language model. arXiv preprint arXiv:2501.18636. https://arxiv.org/abs/2501.18636
SafeRAG is a benchmark that evaluates RAG security against four novel attack tasks (silver noise, inter-context conflict, soft ad, white DoS) that bypass existing retrievers, filters, and LLMs.
[Liang et al. 2025] Liang, Garg, Zilouchian Moghaddam. The SWE-bench illusion: When state-of-the-art llms remember instead of reason. arXiv preprint arXiv:2506.12286. https://arxiv.org/abs/2506.12286
State-of-the-art models name SWE-bench buggy file paths from the issue text alone up to 76
[Libovický and Helcl 2018] Libovický, Helcl. End-to-end non-autoregressive neural machine translation with connectionist temporal classification. https://arxiv.org/abs/1811.04719
This paper proposes an end-to-end non-autoregressive neural machine translation model using CTC, enabling fully parallel decoding without separate multi-step training, evaluated on WMT English-Romanian and English-German.
[Lieber et al. 2024] Lieber, Lenz, Bata, Cohen, Osin, Dalmedigos, Safahi, Meirom, Belinkov, Shalev-Shwartz, Abend, Alon, Asida, Bergman, Glozman, Gokhman, Manevich, Ratner, Rozen, Shwartz, Zusman, Shoham. Jamba: a hybrid transformer-mamba language model. https://arxiv.org/abs/2403.19887
Jamba is a large language model built on a hybrid Transformer-Mamba MoE architecture that achieves 256K token context length and 3x throughput over Mixtral-8x7B while fitting in a single 80GB GPU.
[Lightman et al. 2023] Lightman, Kosaraju, Burda, Edwards, Baker, Lee, Leike, Schulman, Sutskever, Cobbe. Let's verify step by step. arXiv preprint arXiv:2305.20050. https://arxiv.org/abs/2305.20050
Process supervision labels intermediate reasoning steps and outperforms outcome supervision on MATH, producing the PRM800K step-level feedback dataset.
[Lin et al. 2021] Lin, Hilton, Evans. TruthfulQA: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958. https://arxiv.org/abs/2109.07958
TruthfulQA tests whether language models reproduce popular false beliefs, finding that larger imitation-trained models can become less truthful on adversarial questions.
[Lin et al. 2023] Lin, Tang, Tang, Yang, Dang, Han. AWQ: Activation-aware weight quantization for LLM compression and acceleration. https://arxiv.org/abs/2306.00978
AWQ proposes activation-aware per-channel weight scaling for hardware-friendly low-bit (W4A16) LLM quantization, achieving accuracy comparable to mixed-precision FP16 without backpropagation, paired with a TinyChat inference framework delivering over 3x speedup on edge GPUs.
[Lin and others 2025] Lin, others. Continual learning via sparse memory finetuning. arXiv preprint arXiv:2510.15103. https://arxiv.org/abs/2510.15103
Sparse memory finetuning updates only TF-IDF-ranked memory slots in memory-augmented LLMs, enabling continual learning of new facts while avoiding catastrophic forgetting that affects full finetuning and LoRA.
[Lindenbauer et al. 2025] Lindenbauer, Slinko, Felder, Bogomolov, Zharov. The complexity trap: Simple observation masking is as efficient as LLM summarization for agent context management. https://arxiv.org/abs/2508.21433
Simple observation masking halves LLM agent cost on SWE-bench Verified while matching or slightly exceeding the solve rate of LLM-based summarization across five model configurations.
[Lindsey et al. 2025] Lindsey, Gurnee, Ameisen, Chen, Pearce, Turner, Citro, Abrahams, Carter, Hosmer, Marcus, Sklar, Templeton, Bricken, McDougall, Cunningham, Henighan, Jermyn, Jones, Persic, Qi, Thompson, Zimmerman, Rivoire, Conerly, Olah, Batson. On the biology of a large language model. https://transformer-circuits.pub/2025/attribution-graphs/biology.html
This work applies circuit tracing to Claude 3.5 Haiku to investigate the internal mechanisms the model uses across reasoning, poetry planning, multilingual, and arithmetic tasks.
[Linnainmaa 1976] Linnainmaa. Taylor expansion of the accumulated rounding error. BIT Numerical Mathematics. https://doi.org/10.1007/BF01931367
[Lior 2022] Lior. Insuring AI: The role of insurance in artificial intelligence regulation. Harvard Journal of Law & Technology. https://jolt.law.harvard.edu/assets/articlePDFs/v35/Lior-Insuring-AI.pdf
Argues that insurers, by pricing risk and imposing conditions, can act as private regulators of AI the way they historically disciplined fire and automobile safety.
[Lipman et al. 2023] Lipman, Chen, Ben-Hamu, Nickel, Le. Flow matching for generative modeling. https://arxiv.org/abs/2210.02747
Flow Matching (FM) is a simulation-free method for training Continuous Normalizing Flows by regressing conditional vector fields, enabling scalable CNF training with Optimal Transport paths that outperform diffusion models on ImageNet.
[Liu et al. 2023] Liu, Zaharia, Abbeel. Ring attention with blockwise transformers for near-infinite context. https://arxiv.org/abs/2310.01889
Ring Attention distributes long sequences across multiple devices in a ring topology, overlapping key-value block communication with blockwise self-attention computation to enable near-infinite context length without approximations.
[Liu et al. 2023] Liu, Gong, Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. https://arxiv.org/abs/2209.03003
Rectified flow learns an ODE to transport between two distributions by following straight-line paths, enabling high-quality image generation and domain transfer with a single Euler step after iterative reflow.
[Liu et al. 2023] Liu, Li, Wu, Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485. https://arxiv.org/abs/2304.08485
LLaVA connects a CLIP visual encoder to an LLM via a linear projection and applies visual instruction tuning on GPT-4-generated multimodal data to produce a general-purpose vision-language assistant.
[Liu et al. 2023] Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172. https://arxiv.org/abs/2307.03172
Language models show a U-shaped performance curve on long-context tasks, performing best when relevant information is at the start or end of the context and worst when it is in the middle.
[Liu et al. 2024] Liu, Li, Li, Lee. Improved baselines with visual instruction tuning. https://arxiv.org/abs/2310.03744
LLaVA-1.5 shows that replacing LLaVA's linear vision-language connector with an MLP and adding VQA data with response formatting prompts achieves state-of-the-art on 11 multimodal benchmarks using only 1.2M public samples.
[Liu et al. 2024] Liu, Li, Li, Li, Zhang, Shen, Lee. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge. https://llava-vl.github.io/blog/2024-01-30-llava-next/
LLaVA-NeXT improves on LLaVA-1.5 by raising input image resolution up to 4x, enhancing OCR and reasoning via better instruction-tuning data, and matching or exceeding Gemini Pro on several benchmarks while keeping the same minimalist design.
[Liu et al. 2024] Liu, Zhao, Iandola, Lai, Tian, others. MobileLLM: Optimizing sub-billion parameter language models for on-device use cases. https://arxiv.org/abs/2402.14905
MobileLLM shows that deep-and-thin architecture with embedding sharing and GQA outperforms wider sub-billion LLMs, achieving 2.7
[Liu et al. 2024] Liu, Yuan, Jin, Zhong, Xu, Braverman, Chen, Hu. KIVI: a tuning-free asymmetric 2bit quantization for KV cache. https://arxiv.org/abs/2402.02750
KIVI is a tuning-free 2-bit KV cache quantization algorithm that quantizes key cache per-channel and value cache per-token, reducing peak memory by 2.6x and improving throughput by up to 3.47x.
[Liu et al. 2025] Liu, Su, Yao, Jiang, Lai, Du, Qin, Xu, Lu, Yan, others. Muon is scalable for LLM training. arXiv preprint arXiv:2502.16982. https://arxiv.org/abs/2502.16982
Shows the Muon optimizer scales to large LLMs by adding weight decay and tuning the per-parameter update scale, reaching about 2x the compute efficiency of AdamW; trains Moonlight, a 16B MoE model, on 5.7T tokens.
[Liu et al. 2025] Liu, Neubig, Xiong. Midtraining bridges pretraining and posttraining distributions. arXiv preprint arXiv:2510.14865. https://arxiv.org/abs/2510.14865
Defines mid-training as intermediate pretraining-like phases and finds that mixed specialist data can bridge pretraining and post-training distributions, especially for code and math, while reducing forgetting versus pure continued pretraining.
[Liu 2025] Liu. OpenAI could be blowing as much as $15 million per day on silly sora videos. https://www.forbes.com/sites/phoebeliu/2025/11/10/openai-spending-ai-generated-sora-videos/
[Liu et al. 2025] Liu, Diao, Lu, Hu, Dong, Choi, Kautz, Dong. ProRL: Prolonged reinforcement learning expands reasoning boundaries in large language models. arXiv preprint arXiv:2505.24864. https://arxiv.org/abs/2505.24864
ProRL shows that prolonged RL training with KL divergence control and reference policy resets genuinely expands LLM reasoning boundaries, producing Nemotron-Research-Reasoning-Qwen-1.5B which surpasses DeepSeek-R1-7B on several benchmarks.
[Liu et al. 2025] Liu, Chen, Li, Qi, Pang, Du, Lee, Lin. Understanding R1-zero-like training: a critical perspective. arXiv preprint arXiv:2503.20783. https://arxiv.org/abs/2503.20783
Dr. GRPO fixes an optimization bias in GRPO that artificially inflates response length for incorrect outputs, improving token efficiency while maintaining reasoning performance on a 7B model.
[llama.cpp project 2023] llama.cpp project. GGUF file format. https://github.com/ggml-org/llama.cpp
llama.cpp is a C/C++ library for LLM inference on consumer hardware, introducing the GGUF model format and supporting 1.5-bit to 8-bit integer quantization across CPU and GPU backends.
[llm-d project 2025] llm-d project. llm-d: Kubernetes-native distributed inference serving. https://github.com/llm-d/llm-d
llm-d is a CNCF sandbox project founded by Red Hat, Google Cloud, IBM Research, CoreWeave, and NVIDIA that brings prefill/decode disaggregation, prefix-cache-aware routing, and tiered KV-cache offloading to Kubernetes-based LLM serving.
[LMArena 2025] LMArena. Our response to “The Leaderboard Illusion”. https://arena.ai/blog/our-response/
[LMCache team 2025] LMCache team. LMCache: a KV cache management layer for LLM inference. https://github.com/LMCache/LMCache
LMCache is a KV-cache management layer for LLM serving engines that offloads and reuses KV caches across a tiered hierarchy of GPU memory, CPU DRAM, local disk, and remote backends, so a prefix computed once can be served from cache across requests and instances.
[Longpre et al. 2023] Longpre, Mahari, Chen, Obeng-Marnu, Sileo, Brannon, Muennighoff, Khazam, Kabbara, Perisetla, Wu, Shippole, Bollacker, Wu, Villa, Pentland, Hooker. The data provenance initiative: a large scale audit of dataset licensing and attribution in AI. https://arxiv.org/abs/2310.16787
The Data Provenance Initiative audits more than 1,800 text datasets and finds widespread license omissions, license errors, and a divide between commercially open and closed data.
[Longpre et al. 2024] Longpre, Mahari, Lee, Lund, Oderinwale, Brannon, Saxena, Obeng-Marnu, South, Hunter, Klyman, others. Consent in crisis: The rapid decline of the AI data commons. https://arxiv.org/abs/2407.14933
This audit of 14,000 web domains finds fast growth in AI-specific restrictions, with many important C4 sources becoming restricted through robots.txt or terms of service.
[Loshchilov and Hutter 2019] Loshchilov, Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101. https://arxiv.org/abs/1711.05101
Shows L2 regularization and weight decay are not equivalent for Adam, and that decoupling weight decay from the gradient update (AdamW) improves generalization.
[Lou et al. 2024] Lou, Meng, Ermon. Discrete diffusion modeling by estimating the ratios of the data distribution. https://arxiv.org/abs/2310.16834
SEDD introduces score entropy, a novel loss extending score matching to discrete spaces, enabling diffusion language models that beat GPT-2 on perplexity with 32x fewer network evaluations.
[Lovelace et al. 2026] Lovelace, Belardi, Kundurthy, Sudhakar, Weinberger. Prescriptive scaling laws for data constrained training. arXiv preprint arXiv:2605.01640. https://arxiv.org/abs/2605.01640
Adds an explicit overfitting-penalty term to the data-constrained scaling law, finding that larger models overfit repeated data faster, and that strong weight decay sharply reduces the penalty.
[Lu et al. 2022] Lu, Zhou, Bao, Chen, Li, Zhu. DPM-solver: a fast ODE solver for diffusion probabilistic model sampling in around 10 steps. https://arxiv.org/abs/2206.00927
DPM-Solver is a fast, training-free high-order ODE solver for diffusion probabilistic models that generates high-quality samples in 10 to 20 function evaluations, achieving 4 to 16x speedup over prior samplers.
[Luo et al. 2023] Luo, Tan, Huang, Li, Zhao. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv preprint arXiv:2310.04378. https://arxiv.org/abs/2310.04378
Latent Consistency Models (LCMs) distill pre-trained Stable Diffusion into a model that generates 768x768 images in 2 to 4 steps using consistency distillation in latent space.
[Lütke 2025] Lütke. Context engineering over prompt engineering. https://x.com/tobi/status/1935533422589399127
[Lyu et al. 2023] Lyu, Havaldar, Stein, Zhang, Rao, Wong, Apidianaki, Callison-Burch. Faithful chain-of-thought reasoning. arXiv preprint arXiv:2301.13379. https://arxiv.org/abs/2301.13379
Faithful CoT translates natural-language queries into symbolic reasoning chains and uses deterministic solvers, making the executed chain causally responsible for the final answer.
[Ma et al. 2024] Ma, Fang, Wang. DeepCache: Accelerating diffusion models for free. arXiv preprint arXiv:2312.00858. https://arxiv.org/abs/2312.00858
A training-free method that caches high-level U-Net features across adjacent denoising steps and recomputes only the fast-changing parts, 2-4x faster with negligible quality loss.
[Ma et al. 2024] Ma, Wang, Ma, others, Wei. The era of 1-bit llms: All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764. https://arxiv.org/abs/2402.17764
BitNet b1.58 shows that ternary-weight -1, 0, 1 LLMs match FP16 Transformer perplexity and task performance at 3B scale while cutting memory and latency by over 3x.
[Ma et al. 2025] Ma, Pei, Lausen, Karypis. Understanding silent data corruption in LLM training. arXiv preprint arXiv:2502.12340. https://arxiv.org/abs/2502.12340
This paper is the first empirical study of real-world SDC effects on LLM training, showing that corrupted nodes cause parameter drift and loss spikes without triggering explicit failure signals.
[Maclaurin et al. 2015] Maclaurin, Duvenaud, Adams. Autograd: Effortless gradients in numpy. https://github.com/HIPS/autograd
[Madaan et al. 2023] Madaan, Tandon, Gupta, others. Self-refine: Iterative refinement with self-feedback. arXiv preprint arXiv:2303.17651. https://arxiv.org/abs/2303.17651
SELF-REFINE lets a single LLM iteratively generate feedback on its own output and refine it at test time, with no training or RL, improving performance by about 20
[Maini et al. 2024] Maini, Seto, Bai, Grangier, Zhang, Jaitly. Rephrasing the web: a recipe for compute and data-efficient language modeling. https://arxiv.org/abs/2401.16380
WRAP uses an instruction-tuned LLM to rephrase noisy web text into cleaner styles, reducing LLM pre-training compute by  3x and data by  5x compared to training on raw web corpora.
[Malkov and Yashunin 2018] Malkov, Yashunin. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence. https://arxiv.org/abs/1603.09320
HNSW proposes a fully graph-based approximate nearest neighbor search index using a multi-layer proximity graph with logarithmic complexity scaling.
[Manakul et al. 2023] Manakul, Liusie, Gales. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896. https://arxiv.org/abs/2303.08896
SelfCheckGPT detects hallucinations in black-box LLM generations by sampling multiple answers and looking for sentence-level inconsistency.
[Markov et al. 2023] Markov, Zhang, Agarwal, Eloundou, Lee, Adler, Jiang, Weng. A holistic approach to undesired content detection in the real world. AAAI 2023. https://arxiv.org/abs/2208.03274
OpenAI describes a holistic pipeline for real-world content moderation, combining content taxonomy design, active learning, quality-controlled labeling, and synthetic data to detect sexual, hateful, violent, self-harm, and harassment content.
[Martucci 2026] Martucci. GE vernova gas turbine backlog hits 100 GW as prices rise. https://www.utilitydive.com/news/ge-vernova-gas-turbine-backlog-hits-100-gw-as-prices-rise/818332/
GE Vernova's gas turbine backlog plus reservations reached 100 GW in Q1 2026, with the company expecting reservations sold out through 2030 by year-end.
[Mazeika et al. 2024] Mazeika, Phan, Yin, Zou, Wang, Mu, Sakhaee, Li, Basart, Li, Forsyth, Hendrycks. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249. https://arxiv.org/abs/2402.04249
HarmBench is a standardized benchmark with 510 harmful behaviors and an evaluation pipeline that compares 18 automated red-teaming methods across 33 LLMs to enable rigorous, reproducible attack-defense co-development.
[McCandlish et al. 2018] McCandlish, Kaplan, Amodei, OpenAI Dota Team. An empirical model of large-batch training. arXiv preprint arXiv:1812.06162. https://arxiv.org/abs/1812.06162
Introduces the gradient noise scale to predict the critical batch size, beyond which larger batches stop reducing the number of training steps and waste compute.
[McCulloch and Pitts 1943] McCulloch, Pitts. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics.
McCulloch and Pitts give the first mathematical model of a neuron, showing networks of threshold logic units can compute any logical function.
[McKinsey & Company 2025] McKinsey & Company. The automation curve in agentic commerce. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-automation-curve-in-agentic-commerce
[McMahan et al. 2013] McMahan, Holt, Sculley, Young, Ebner, Grady, Nie, Phillips, Davydov, Golovin, Chikkerur, Liu, Wattenberg, Hrafnkelsson, Boulos, Kubica. Ad click prediction: a view from the trenches. https://research.google/pubs/ad-click-prediction-a-view-from-the-trenches/
Google's report on its production ad click-through predictor: FTRL online learning over billions of sparse features, plus the calibration, feature management, and monitoring discipline that kept it operable.
[McNemar 1947] McNemar. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika. https://doi.org/10.1007/BF02295996
McNemar's test compares paired binary outcomes by focusing on examples where two systems disagree, which is often the right design for benchmark A/B comparisons on the same items.
[Mehta et al. 2024] Mehta, Sekhavat, Cao, Rastegari, others. OpenELM: An efficient language model family with open training and inference framework. arXiv preprint arXiv:2404.14619. https://arxiv.org/abs/2404.14619
OpenELM introduces layer-wise scaling in a decoder-only transformer to allocate parameters non-uniformly across layers, achieving 2.36
[Mei et al. 2025] Mei, Yao, Ge, Wang, Bi, Cai, Liu, Li, Li, Zhang, Zhou, Zhang, Yan, Wang, Pan, Yan, others. A survey of context engineering for large language models. arXiv preprint arXiv:2507.13334. https://arxiv.org/abs/2507.13334
A survey of context engineering for LLMs, providing a unified taxonomy covering RAG, memory systems, tool-integrated reasoning, and multi-agent systems across 1400+ papers.
[Mem0 2025] Mem0. Introducing OpenMemory MCP. https://mem0.ai/blog/introducing-openmemory-mcp
[Meng et al. 2022] Meng, Bau, Andonian, Belinkov. Locating and editing factual associations in GPT. https://arxiv.org/abs/2202.05262
ROME uses causal mediation analysis to locate factual associations in mid-layer feed-forward modules of GPT and introduces a rank-one weight editing method to update specific facts.
[Meng et al. 2023] Meng, Sharma, Andonian, Belinkov, Bau. Mass-editing memory in a transformer. https://arxiv.org/abs/2210.07229
MEMIT scales knowledge editing in large language models to thousands of simultaneous fact updates by distributing parameter changes across a range of critical MLP layers.
[Meng et al. 2024] Meng, Xia, Chen. SimPO: Simple preference optimization with a reference-free reward. https://arxiv.org/abs/2405.14734
SimPO replaces DPO's reference-model reward with a length-normalized average log probability and a target reward margin, eliminating the reference model while outperforming DPO by up to 7.5 points on Arena-Hard.
[Meta 2025] Meta. Improving your recommendations on our apps with AI at meta. https://about.fb.com/news/2025/10/improving-your-recommendations-apps-ai-meta/
[Meta 2025] Meta. Private processing for WhatsApp. https://ai.meta.com/static-resource/private-processing-technical-whitepaper
Meta's technical whitepaper rebuilding the Private Cloud Compute requirements on commodity parts: SEV-SNP confidential VMs, NVIDIA GPU confidential mode, and anonymous relays, at WhatsApp scale.
[Meta 2025] Meta. Collective communication for 100k+ gpus. arXiv preprint arXiv:2510.20171. https://arxiv.org/abs/2510.20171
NCCLX is Meta's collective communication framework built on NCCL for Llama4, supporting 100K+ GPUs with zero-copy SM-free transfers, fault-tolerant AllReduce, and GPU-resident collectives for MoE inference.
[Meta 2025] Meta. Meta and blue owl capital to develop hyperion data center. https://about.fb.com/news/2025/10/meta-blue-owl-capital-develop-hyperion-data-center/
A  $27B joint venture in which a credit fund owns most of a datacenter campus and Meta leases the capacity back, keeping the debt off Meta's own balance sheet.
[Meta AI 2025] Meta AI. The llama 4 herd: The beginning of a new era of natively multimodal AI innovation. https://ai.meta.com/blog/llama-4-multimodal-intelligence/
[Meta AI 2025] Meta AI. Llama guard 4 model card (llama-guard-4-12B). https://huggingface.co/meta-llama/Llama-Guard-4-12B
[METR 2025] METR. Details about metr's preliminary evaluation of GPT-4.5. https://metr.org/blog/2025-02-27-gpt-4-5-evals/
A pre-deployment autonomy evaluation with a candid account of its own limits: short lab-controlled access to a checkpoint, described by the evaluator as a first step rather than an audit.
[METR 2025] METR. Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv preprint arXiv:2507.09089. https://arxiv.org/abs/2507.09089
A randomized controlled trial with 16 experienced open-source developers found that early-2025 AI tools (Cursor Pro with Claude 3.5/3.7 Sonnet) increased task completion time by 19
[METR 2026] METR. Time horizon 1.1. https://metr.org/blog/2026-1-29-time-horizon-1-1/
METR releases Time Horizon 1.1, an updated benchmark measuring AI agent task-completion time horizons using more tasks and a new eval infrastructure.
[METR 2026] METR. A simpler AI timelines model predicts 99% AI R&D automation in  2032. https://metr.org/notes/2026-02-10-simpler-ai-timelines-model/
METR presents an 8-parameter model for forecasting AI timelines, predicting roughly 99
[Mialon et al. 2023] Mialon, Fourrier, Swift, Wolf, LeCun, Scialom. GAIA: a benchmark for general AI assistants. arXiv preprint arXiv:2311.12983. https://arxiv.org/abs/2311.12983
GAIA is a benchmark of 466 real-world questions for general AI assistants, where humans score 92
[Michaud et al. 2023] Michaud, Liu, Girit, Tegmark. The quantization model of neural scaling. https://arxiv.org/abs/2303.13506
Proposes that skills are discrete quanta learned in order of decreasing use frequency; when those frequencies follow a power law, learning them in order produces smooth power-law loss and recasts emergence as quanta switching on.
[Micikevicius et al. 2017] Micikevicius, Narang, Alben, Diamos, Elsen, Garcia, Ginsburg, Houston, Kuchaiev, Venkatesh, Wu. Mixed precision training. https://arxiv.org/abs/1710.03740
This paper presents mixed precision training, combining FP16 storage and arithmetic with FP32 master weights, loss-scaling, and FP32 accumulation to halve memory use without accuracy loss.
[Micikevicius et al. 2022] Micikevicius, Stosic, Burgess, Cornea, Dubey, Grisenthwaite, Ha, Heinecke, Judd, Kamalu, Mellempudi, Oberman, Shoeybi, Siu, Wu. FP8 formats for deep learning. https://arxiv.org/abs/2209.05433
This paper proposes an FP8 binary interchange format with two encodings, E4M3 and E5M2, and demonstrates that FP8 training matches 16-bit quality across CNNs, RNNs, and Transformers up to 175B parameters.
[Microsoft 2024] Microsoft. MLOps and GenAIOps for AI workloads on azure. https://learn.microsoft.com/en-us/azure/well-architected/ai/mlops-genaiops
Microsoft's workload guidance treats AI operations as lifecycle management for nondeterministic systems, including monitoring, drift, deployment, governance, and automation.
[Microsoft 2025] Microsoft. The next chapter of the Microsoft-OpenAI partnership. https://blogs.microsoft.com/blog/2025/10/28/the-next-chapter-of-the-microsoft-openai-partnership/
Microsoft holds roughly 27
[Milakov and Gimelshein 2018] Milakov, Gimelshein. Online normalizer calculation for softmax. arXiv preprint arXiv:1805.02867. https://arxiv.org/abs/1805.02867
A single-pass softmax that maintains a running maximum and normalizer, the four-page trick that years later made tiled exact attention possible.
[Miller 2024] Miller. Adding error bars to evals: a statistical approach to language model evaluations. arXiv preprint arXiv:2411.00640. https://arxiv.org/abs/2411.00640
An Anthropic treatment of LLM evaluation as statistical inference: report standard errors, use clustered standard errors when questions come in groups, analyze paired differences between models, and plan sample sizes with power analysis.
[Min et al. 2023] Min, Krishna, Lyu, Lewis, Yih, Koh, Iyyer, Zettlemoyer, Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251. https://arxiv.org/abs/2305.14251
FActScore decomposes long-form generations into atomic facts and scores the percentage supported by reliable sources, making mixed factual and non-factual passages measurable.
[MiniMax 2025] MiniMax. MiniMax-01: Scaling foundation models with lightning attention. arXiv preprint arXiv:2501.08313. https://arxiv.org/abs/2501.08313
MiniMax-01 interleaves lightning attention (a linear attention) with softmax attention in a 456B-parameter MoE model, reaching million-token contexts at frontier quality.
[Mitchell et al. 2019] Mitchell, Wu, Zaldivar, Barnes, Vasserman, Hutchinson, Spitzer, Raji, Gebru. Model cards for model reporting. https://arxiv.org/abs/1810.03993
Model cards document intended uses, evaluation procedures, and disaggregated performance to support transparent decisions about model deployment.
[Mithril Security 2023] Mithril Security. PoisonGPT: How we hid a lobotomized LLM on hugging face to spread fake news. https://mithrilsecurity.io/blog/poisongpt
A demonstration editing specific false facts into a model's weights and uploading it under an org name one letter off from a well-known lab, combining a weight edit with typosquatting.
[mlc-ai 2023] mlc-ai. MLC-LLM: Universal LLM deployment engine with ML compilation. https://github.com/mlc-ai/mlc-llm
MLC-LLM is an open-source universal LLM deployment engine that uses ML compilation to run large language models across diverse hardware backends.
[Model Context Protocol 2025] Model Context Protocol. MCP transport future: Streamable HTTP replaces SSE. https://blog.modelcontextprotocol.io/posts/2025-12-19-mcp-transport-future/
The MCP Transport Working Group outlines a roadmap to make MCP stateless, elevate sessions to the data model layer, and introduce Server Cards to scale remote deployments.
[Model Context Protocol 2025] Model Context Protocol. Transports: Mcp-Session-Id and Last-Event-ID. https://modelcontextprotocol.io/specification/2025-11-25/basic/transports
The MCP specification's transport layer defines two mechanisms for JSON-RPC message exchange: stdio for subprocess communication and Streamable HTTP (replacing SSE) for networked MCP servers.
[Model Context Protocol 2025] Model Context Protocol. Key changes: MCP specification 2025-06-18. https://modelcontextprotocol.io/specification/2025-06-18/changelog
The 2025-06-18 MCP revision classifies servers as OAuth resource servers, requires clients to implement RFC 8707 resource indicators, removes JSON-RPC batching, and adds structured tool output and elicitation.
[Model Context Protocol 2025] Model Context Protocol. Model context protocol specification, revision 2025-06-18: Authorization. https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization
[Model Context Protocol 2026] Model Context Protocol. The 2026 MCP roadmap. https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/
The official MCP 2026 roadmap prioritizes transport scalability, agent communication, governance delegation, and enterprise readiness as the four areas driving protocol development.
[Model Context Protocol 2026] Model Context Protocol. The 2026-07-28 MCP specification release candidate. https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/
The 2026-07-28 MCP release candidate makes the protocol stateless: SEP-2567 removes the Mcp-Session-Id header and protocol-level sessions, SEP-2575 removes the initialize handshake, and servers carry state via explicit handles.
[Mohan et al. 2021] Mohan, Phanishayee, Chidambaram. CheckFreq: Frequent, fine-grained DNN checkpointing. https://www.usenix.org/conference/fast21/presentation/mohan
[Moonshot AI 2025] Moonshot AI. Kimi-researcher: End-to-end RL training for emerging agentic capabilities. https://moonshotai.github.io/Kimi-Researcher/
Kimi-Researcher is trained entirely through end-to-end agentic reinforcement learning, averaging 23 reasoning steps and exploring over 200 URLs per task, reaching 26.9
[Morgan Stanley Research 2025] Morgan Stanley Research. Bridging a $1.5tr data center financing gap. https://www.morganstanley.com/content/dam/msdotcom/en/assets/pdfs/Research_Bridging-Data-Center-Gap.pdf
Estimates roughly $2.9T of global datacenter spend through 2028 against  $1.4T of hyperscaler cash flow, leaving a  $1.5T financing gap,  $800B of it from asset-backed private credit.
[Motamed et al. 2025] Motamed, Culp, others. Do generative video models understand physical principles?. arXiv preprint arXiv:2501.09038. https://arxiv.org/abs/2501.09038
Physics-IQ is a real-world benchmark of 396 videos covering five physics domains that reveals current generative video models score at most 29.5
[Muennighoff et al. 2023] Muennighoff, Rush, Barak, Le Scao, Piktus, Tazi, Pyysalo, Wolf, Raffel. Scaling data-constrained language models. arXiv preprint arXiv:2305.16264. https://arxiv.org/abs/2305.16264
Studies training under limited data and finds that repeating data for up to about four epochs is nearly as good as fresh data, then proposes a scaling law for the decaying value of repeated tokens and excess parameters.
[Muennighoff et al. 2023] Muennighoff, Tazi, Magne, Reimers. MTEB: Massive text embedding benchmark. https://arxiv.org/abs/2210.07316
MTEB introduces a benchmark spanning 8 embedding tasks, 58 datasets, and 112 languages, finding that no single text embedding method dominates all tasks.
[Muennighoff et al. 2025] Muennighoff, Yang, Shi, Li, Fei-Fei, Hajishirzi, Zettlemoyer, Liang, Candès, Hashimoto. s1: Simple test-time scaling. arXiv preprint arXiv:2501.19393. https://arxiv.org/abs/2501.19393
s1 achieves test-time scaling by SFT on 1,000 curated reasoning samples and a budget forcing technique, yielding s1-32B that outperforms o1-preview on competition math.
[Nakano et al. 2021] Nakano, Hilton, Balaji, Wu, Ouyang, Kim, Hesse, Jain, Kosaraju, Saunders, others. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332. https://arxiv.org/abs/2112.09332
WebGPT fine-tunes GPT-3 to answer long-form questions by browsing the web via a text-based environment, trained with imitation learning and optimized with human feedback using reward modeling and rejection sampling.
[Narayanan et al. 2019] Narayanan, Harlap, Phanishayee, Seshadri, Devanur, Ganger, Gibbons, Zaharia. PipeDream: Generalized pipeline parallelism for DNN training. https://doi.org/10.1145/3341301.3359646
PipeDream combines pipeline parallelism with data parallelism to reduce inter-GPU communication by up to 95
[Narayanan et al. 2021] Narayanan, Shoeybi, Casper, LeGresley, Patwary, Korthikanti, Vainbrand, Kashinkunti, Bernauer, Catanzaro, Phanishayee, Zaharia. Efficient large-scale language model training on GPU clusters using megatron-LM. https://arxiv.org/abs/2104.04473
Megatron-LM combines tensor parallelism, pipeline parallelism, and data parallelism (PTD-P) with a novel interleaved pipeline schedule to train trillion-parameter language models at 502 petaFLOP/s on 3072 GPUs.
[National Institute of Standards and Technology 2023] National Institute of Standards and Technology. Artificial intelligence risk management framework (AI RMF 1.0). https://www.nist.gov/itl/ai-risk-management-framework
NIST AI RMF 1.0 organizes AI risk management into govern, map, measure, and manage functions across the AI lifecycle.
[National Institute of Standards and Technology 2024] National Institute of Standards and Technology. Artificial intelligence risk management framework: Generative artificial intelligence profile. https://doi.org/10.6028/NIST.AI.600-1
The NIST generative AI profile adapts the AI RMF to risks specific to generative systems, which this chapter translates into operating controls.
[Naumov et al. 2019] Naumov, Mudigere, Shi, Huang, Sundaraman, Park, Wang, Gupta, Wu, Azzolini, Dzhulgakov, Mallevich, Cherniavskii, Lu, Krishnamoorthi, Yu, Kondratenko, Pereira, Chen, Chen, Rao, Jia, Xiong, Smelyanskiy. Deep learning recommendation model for personalization and recommendation systems. arXiv preprint arXiv:1906.00091. https://arxiv.org/abs/1906.00091
Meta's reference recommendation architecture: terabyte-scale sparse embedding tables with a small dense network on top, the workload shape that dominated pre-LLM AI datacenters.
[Ni et al. 2022] Ni, Hernández Ábrego, Constant, Ma, Hall, Cer, Yang. Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models. https://arxiv.org/abs/2108.08877
Sentence-T5 (ST5) extracts sentence embeddings from T5 encoder-decoder models using three methods, outperforming SentenceBERT and SimCSE on STS and transfer benchmarks, with scaling to 11B parameters yielding consistent gains.
[Ni et al. 2022] Ni, Qu, Lu, Dai, Hernández Ábrego, Ma, Zhao, Luan, Hall, Chang, Yang. Large dual encoders are generalizable retrievers. https://arxiv.org/abs/2112.07899
GTR shows that scaling up T5-based dual encoders to 5B parameters, with fixed bottleneck embedding size and multi-stage training, substantially improves out-of-domain retrieval generalization on BEIR.
[Nichol and Dhariwal 2021] Nichol, Dhariwal. Improved denoising diffusion probabilistic models. https://arxiv.org/abs/2102.09672
This paper shows that DDPMs achieve competitive log-likelihoods via learned reverse-process variances and a hybrid objective, enabling high-quality sampling with 20x fewer forward passes.
[Nickolls et al. 2008] Nickolls, Buck, Garland, Skadron. Scalable parallel programming with CUDA. ACM Queue. https://queue.acm.org/detail.cfm?id=1365500
The canonical description of CUDA's execution model, grids, blocks, warps, and SIMT, written by its architects the year after launch.
[Nie et al. 2025] Nie, Zhu, You, Zhang, Ou, Hu, Zhou, Lin, Wen, Li. Large language diffusion models. arXiv preprint arXiv:2502.09992. https://arxiv.org/abs/2502.09992
LLaDA introduces a masked diffusion LLM trained from scratch via pre-training and SFT, showing that scalability, in-context learning, and instruction-following do not require autoregressive modeling.
[Northcutt et al. 2021] Northcutt, Athalye, Mueller. Pervasive label errors in test sets destabilize machine learning benchmarks. NeurIPS Datasets and Benchmarks Track. https://arxiv.org/abs/2103.14749
This paper identifies pervasive label errors averaging 3.3
[Noukhovitch et al. 2024] Noukhovitch, Huang, Xhonneux, Hosseini, Agarwal, Courville. Asynchronous RLHF: Faster and more efficient off-policy RL for language models. arXiv preprint arXiv:2410.18252. https://arxiv.org/abs/2410.18252
Asynchronous RLHF decouples generation and training onto separate GPUs, using off-policy Online DPO to achieve 40-70
[Novikov et al. 2025] Novikov, Vu, Eisenberger, Dupont, Huang, Wagner, Shirobokov, Kozlovskii, Ruiz, Mehrabian, Kumar, See, Chaudhuri, Holland, Davies, Nowozin, Kohli, Balog. AlphaEvolve: a coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131. https://arxiv.org/abs/2506.13131
AlphaEvolve uses an evolutionary coding-agent loop with evaluator feedback to improve algorithms and infrastructure code, illustrating discovery when evaluation is engineered.
[NucNet 2025] NucNet. Constellation secures $1 billion federal loan for three mile island restart. https://www.nucnet.org/news/constellation-secures-usd1-billion-federal-loann-for-three-mile-island-restart-11-3-2025
[NVIDIA 2023] NVIDIA. TensorRT-LLM. https://github.com/NVIDIA/TensorRT-LLM
TensorRT-LLM is NVIDIA's open-source library providing a Python API and state-of-the-art optimizations for efficient LLM inference on NVIDIA GPUs, plus C++/Python runtimes for orchestrating inference execution.
[NVIDIA 2024] NVIDIA. NCCL: NVIDIA collective communications library. https://github.com/NVIDIA/nccl
NCCL is NVIDIA's open-source library providing optimized collective communication primitives for multi-GPU systems.
[NVIDIA 2024] NVIDIA. Nemotron-4 340B technical report. https://arxiv.org/abs/2406.11704
NVIDIA releases Nemotron-4 340B, a family of open-access base, instruct, and reward models where over 98
[NVIDIA 2025] NVIDIA. Nemotron-h: a family of accurate and efficient hybrid mamba-transformer models. arXiv preprint arXiv:2504.03624. https://arxiv.org/abs/2504.03624
Nemotron-H replaces most attention layers with Mamba layers in 8B and 56B models, matching Transformer accuracy at up to 3x faster long-context inference.
[NVIDIA 2025] NVIDIA. Pretraining large language models with NVFP4. arXiv preprint arXiv:2509.25149. https://arxiv.org/abs/2509.25149
Trains a 12B model over 10 trillion tokens in the NVFP4 4-bit microscaling format, using Random Hadamard transforms, two-dimensional scaling, and stochastic rounding to match an fp8 baseline.
[NVIDIA 2025] NVIDIA. NVIDIA Dynamo: a datacenter-scale distributed inference serving framework. https://github.com/ai-dynamo/dynamo
NVIDIA Dynamo is an open-source datacenter-scale inference serving framework built around prefill/decode disaggregation, KV-cache-aware request routing, and tiered KV-cache offloading across GPU memory, host DRAM, SSD, and networked storage.
[NVIDIA 2025] NVIDIA. OpenAI and NVIDIA announce strategic partnership to deploy 10GW of NVIDIA systems. https://nvidianews.nvidia.com/news/openai-and-nvidia-announce-strategic-partnership-to-deploy-10gw-of-nvidia-systems
A letter of intent to deploy at least 10GW of NVIDIA systems, with NVIDIA intending to invest up to $100B in OpenAI progressively as each gigawatt is deployed.
[Nygard 2018] Nygard. Release it! Design and deploy production-ready software. Pragmatic Bookshelf.
Nygard catalogs stability and resilience patterns (circuit breakers, bulkheads, timeouts) for designing production software that survives failure.
[Odlyzko 2010] Odlyzko. Collective hallucinations and inefficient markets: The british railway mania of the 1840s. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1537338
The railway mania as the base case for how a genuinely transformative technology can coexist with catastrophic capital destruction, with demand evidence ignored in plain sight.
[Office of Governor Kathy Hochul 2025] Office of Governor Kathy Hochul. Governor hochul signs nation-leading legislation to require AI frameworks for AI frontier models. https://www.governor.ny.gov/news/governor-hochul-signs-nation-leading-legislation-require-ai-frameworks-ai-frontier-models
[Ogenrwot and Businge 2026] Ogenrwot, Businge. AgenticFlict: a large-scale dataset of merge conflicts in AI coding agent pull requests on GitHub. https://arxiv.org/abs/2604.03551
AGENTICFLICT is a large-scale dataset of 336K+ merge conflict regions extracted from 142K+ AI coding agent pull requests on GitHub, revealing a 27.67
[Okta 2025] Okta. Okta for AI agents, early access. https://www.okta.com/blog/ai/okta-ai-agents-early-access-announcement/
Okta for AI Agents (Early Access) provides AI agents with a dedicated identity layer covering visibility, lifecycle management, and governance.
[Olah et al. 2020] Olah, Cammarata, Schubert, Goh, Petrov, Carter. Zoom in: An introduction to circuits. Distill. https://distill.pub/2020/circuits/zoom-in/
This Distill article proposes that neural networks contain interpretable "circuits" of neurons encoding meaningful algorithms, and advances three hypotheses: features, circuits, and universality across models.
[OLMo Team 2025] OLMo Team. 2 olmo 2 furious. https://arxiv.org/abs/2501.00656
OLMo 2 is a fully open family of 7B, 13B, and 32B dense autoregressive language models that release weights, training data, code, and recipes, matching open-weight models like Llama 3.1 at the performance-to-FLOPs Pareto frontier.
[Olsson et al. 2022] Olsson, Elhage, Nanda, Joseph, DasSarma, Henighan, Mann, Askell, Bai, Chen, Conerly, Drain, Ganguli, Hatfield-Dodds, Hernandez, Johnston, Jones, Kernion, Lovitt, Ndousse, Amodei, Brown, Clark, Kaplan, McCandlish, Olah. In-context learning and induction heads. https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
This paper identifies induction heads, a two-attention-head circuit in transformers, as the primary mechanistic source of in-context learning across model sizes.
[Olston et al. 2017] Olston, Fiedel, Gorovoy, Harmsen, Lao, Li, Rajashekhar, Ramesh, Soyke. TensorFlow-serving: Flexible, high-performance ML serving. arXiv preprint arXiv:1712.06139. https://arxiv.org/abs/1712.06139
[Oord et al. 2018] Oord, Li, Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748. https://arxiv.org/abs/1807.03748
CPC learns representations from unlabeled data by predicting future observations in latent space using a contrastive loss (InfoNCE), demonstrating strong results across speech, images, text, and reinforcement learning.
[Open Source Initiative 2024] Open Source Initiative. The open source AI definition 1.0. https://opensource.org/ai/open-source-ai-definition
The OSI definition treats Open Source AI as requiring use, study, modification, and sharing freedoms, with data information, code, and parameters available in the preferred form for modification.
[Open X-Embodiment Collaboration 2023] Open X-Embodiment Collaboration. Open x-embodiment: Robotic learning datasets and RT-x models. arXiv preprint arXiv:2310.08864. https://arxiv.org/abs/2310.08864
Open X-Embodiment assembles a dataset of 1M+ trajectories from 22 robot embodiments across 21 institutions and trains RT-X models that show positive transfer across robot platforms.
[OpenAI 2020] OpenAI. OpenAI standardizes on PyTorch. https://openai.com/index/openai-pytorch/
[OpenAI 2023] OpenAI. GPT-4 technical report. https://cdn.openai.com/papers/gpt-4.pdf
GPT-4 is a multimodal Transformer trained with RLHF that reaches human-level performance on professional and academic benchmarks, with capabilities predictable via scaling laws from small runs using 1,000x less compute.
[OpenAI 2024] OpenAI. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/
[OpenAI 2024] OpenAI. Learning to reason with llms. https://openai.com/index/learning-to-reason-with-llms/
[OpenAI 2024] OpenAI. Memory and new controls for ChatGPT. https://openai.com/index/memory-and-new-controls-for-chatgpt/
The first consumer deployment of assistant memory: user-manageable saved memories, deletion controls, and temporary chats that leave no trace.
[OpenAI 2024] OpenAI. Introducing SWE-bench verified. https://openai.com/index/introducing-swe-bench-verified/
[OpenAI 2025] OpenAI. Sycophancy in GPT-4o: What happened and what we're doing about it. https://openai.com/index/sycophancy-in-gpt-4o/
The post-mortem of the sycophantic model update, including the admission that user memory contributed to exacerbating sycophancy's effects in some cases.
[OpenAI 2025] OpenAI. Introducing gpt-realtime and Realtime API updates for production voice agents. https://openai.com/index/introducing-gpt-realtime/
[OpenAI 2025] OpenAI. Introducing 4o image generation. https://openai.com/index/introducing-4o-image-generation/
[OpenAI 2025] OpenAI. Sora 2 is here. https://openai.com/index/sora-2/
[OpenAI 2025] OpenAI. Model spec. https://model-spec.openai.com/2025-12-18.html
The Model Spec defines intended model behavior and authority levels for resolving conflicting instructions.
[OpenAI 2025] OpenAI. Introducing GPT-5. https://openai.com/index/introducing-gpt-5/
[OpenAI 2025] OpenAI. Introducing deep research. https://openai.com/index/introducing-deep-research/
OpenAI's Deep Research agent was trained end to end with reinforcement learning on real-world browsing and Python tool-use tasks, using the same RL methods behind the o1 reasoning model.
[OpenAI 2025] OpenAI. Introducing operator. https://openai.com/index/introducing-operator/
The Computer-Using Agent: screenshots in, mouse and keyboard actions out, running in a hosted remote browser, and roughly tripling OSWorld's best score within a year of the benchmark's release.
[OpenAI 2025] OpenAI. GDPval: Evaluating AI model performance on real-world economically valuable tasks. arXiv preprint arXiv:2510.04374. https://arxiv.org/abs/2510.04374
GDPval is a benchmark of 1,320 economically grounded tasks across 44 occupations and 9 U.S. GDP sectors, showing frontier models approaching expert parity via pairwise human comparison.
[OpenAI 2025] OpenAI. Introducing gpt-oss. https://openai.com/index/introducing-gpt-oss/
OpenAI releases gpt-oss-120b and gpt-oss-20b, two open-weight mixture-of-experts reasoning models under Apache 2.0, its first open-weight language models since GPT-2, with the larger model near o4-mini on reasoning benchmarks while fitting a single 80 GB GPU.
[OpenAI 2026] OpenAI. Codex CLI: Conversation and code-reverting rewind. https://github.com/openai/codex/issues/11626
A GitHub issue on openai/codex CLI requesting a /rewind command that atomically restores both conversation history and Codex-applied file edits to a selected checkpoint.
[OpenAI 2026] OpenAI. Dreaming: Better memory for a more helpful ChatGPT. https://openai.com/index/chatgpt-memory-dreaming/
[OpenID Foundation 2026] OpenID Foundation. AuthZEN authorization API 1.0. https://openid.net/specs/authorization-api-1_0.html
[OpenSSF AI/ML Working Group 2025] OpenSSF AI/ML Working Group. Launch of model signing v1.0. https://openssf.org/blog/2025/04/04/launch-of-model-signing-v1-0/
A signing library and format for models of any format and size, verifying against a signer identity with a transparency log, answering provenance rather than safety.
[OpenTelemetry n.d.] OpenTelemetry. OpenTelemetry GenAI conventions. https://github.com/open-telemetry/semantic-conventions-genai
OpenTelemetry GenAI semantic conventions specify attributes and spans for tracing model calls, agent steps, and GenAI system behavior across vendors.
[Ou et al. 2025] Ou, Nie, Xue, Zhu, Sun, Li, Li. Your absorbing discrete diffusion secretly models the conditional distributions of clean data. https://arxiv.org/abs/2406.03736
RADD reparameterizes the concrete score of absorbing discrete diffusion as time-independent conditional distributions of clean data, enabling output caching for faster sampling and achieving SOTA perplexity among diffusion models at GPT-2 scale.
[Ou 2026] Ou. Goalless agents. https://changkun.de/blog/posts/goalless-agents
Empirical experiments show that goalless AI agents (Claude and Codex) exhibit stable training-imprinted priors, and an 80/20 exploration-exploitation rhythm produces directed depth that neither rigid pipelines nor unconstrained freedom achieves.
[Ouyang et al. 2022] Ouyang, Wu, Jiang, Almeida, Wainwright, Mishkin, Zhang, Agarwal, Slama, Ray, Schulman, Hilton, Kelton, Miller, Simens, Askell, Welinder, Christiano, Leike, Lowe. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2203.02155
InstructGPT fine-tunes GPT-3 with RLHF via SFT, reward modeling, and PPO to align language model outputs with human intent, showing a 1.3B model preferred over 175B GPT-3.
[Ouyang et al. 2022] Ouyang, Wu, Jiang, Almeida, Wainwright, Mishkin, Zhang, Agarwal, Slama, Ray, Schulman, Hilton, Kelton, Miller, Simens, Askell, Welinder, Christiano, Leike, Lowe. Training language models to follow instructions with human feedback. https://arxiv.org/abs/2203.02155
InstructGPT fine-tunes GPT-3 with RLHF (SFT then PPO against a human preference reward model) to follow user instructions, with a 1.3B model preferred over the 175B GPT-3 baseline.
[Ouyang et al. 2025] Ouyang, Guo, Arora, Zhang, Hu, Ré, Mirhoseini. KernelBench: Can llms write efficient GPU kernels?. https://arxiv.org/abs/2502.10517
The benchmark asking whether models can write correct, faster-than-PyTorch GPU kernels: 250 workloads, a speedup-aware metric, and frontier models beating the baseline in under 20
[OWASP CycloneDX 2024] OWASP CycloneDX. CycloneDX v1.6 with AI/ML bill of materials (ML-BOM). https://cyclonedx.org/capabilities/mlbom/
The ML extension of a software bill-of-materials standard, letting a model ship a machine-readable inventory of its components, datasets, and lineage.
[Packer et al. 2023] Packer, Wooders, Lin, Fang, Patil, Stoica, Gonzalez. MemGPT: Towards llms as operating systems. https://arxiv.org/abs/2310.08560
MemGPT introduces virtual context management for LLMs, using an OS-inspired hierarchical memory system to page data between a fixed context window and external storage, enabling unbounded context for document analysis and multi-session chat.
[Pagnoni et al. 2024] Pagnoni, Pasunuru, Rodriguez, Nguyen, Muller, others. Byte latent transformer: Patches scale better than tokens. arXiv preprint arXiv:2412.09871. https://arxiv.org/abs/2412.09871
BLT drops the tokenizer and groups raw bytes into dynamically sized patches based on next-byte entropy, matching tokenization-based models in FLOP-controlled scaling up to 8B parameters.
[Pai 2025] Pai. Designing large language model applications: a holistic approach. O'Reilly Media. https://www.oreilly.com/library/view/designing-large-language/9781098150495/
An O'Reilly book by Suhas Pai covering practical design patterns and engineering decisions for building production LLM applications in enterprises.
[Pan et al. 2024] Pan, Wang, Neubig, Jaitly, Ji, Suhr, Zhang. Training software engineering agents and verifiers with SWE-gym. arXiv preprint arXiv:2412.21139. https://arxiv.org/abs/2412.21139
SWE-Gym introduces the first training environment with 2,438 real-world Python GitHub issues, executable environments, and unit tests to train and evaluate SWE agents via supervised fine-tuning and inference-time scaling.
[Panickssery et al. 2024] Panickssery, Bowman, Feng. LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076. https://arxiv.org/abs/2404.13076
LLMs such as GPT-4 and Llama 2 exhibit self-preference bias in evaluation, and this bias is linearly correlated with the model's ability to recognize its own outputs.
[Parasuraman et al. 2000] Parasuraman, Sheridan, Wickens. A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans. https://doi.org/10.1109/3468.844354
Parasuraman, Sheridan, and Wickens separate automation by information acquisition, analysis, decision selection, and action implementation, each with distinct levels of human involvement.
[Parikh and Surapaneni 2025] Parikh, Surapaneni. Announcing agent payments protocol (AP2). https://cloud.google.com/blog/products/ai-machine-learning/announcing-agents-to-payments-ap2-protocol
Google's payment protocol with 60+ partners: signed intent and cart mandates form a tamper-evident chain answering who authorized what, whether the request reflects the user's intent, and who is accountable.
[Park et al. 2023] Park, O'Brien, Cai, Morris, Liang, Bernstein. Generative agents: Interactive simulacra of human behavior. arXiv preprint arXiv:2304.03442. https://arxiv.org/abs/2304.03442
The memory stream with retrieval scored by recency, importance, and relevance, plus periodic reflection that synthesizes higher-level memories: the conceptual template every later memory system echoes.
[Paszke et al. 2019] Paszke, Gross, Massa, Lerer, Bradbury, Chanan, Killeen, Lin, Gimelshein, Antiga, Desmaison, Köpf, Yang, DeVito, Raison, Tejani, Chilamkurthy, Steiner, Fang, Bai, Chintala. PyTorch: An imperative style, high-performance deep learning library. https://arxiv.org/abs/1912.01703
PyTorch's design paper: an eager tape-based autodiff with the hot path in C++, a caching CUDA allocator, and the argument for why usability and performance are not opposites.
[Patwardhan et al. 2025] Patwardhan, Dias, Proehl, others. GDPval: Evaluating AI model performance on real-world economically valuable tasks. https://arxiv.org/abs/2510.04374
GDPval evaluates model output on 1,320 real work deliverables from 44 occupations in the top nine US GDP sectors, finding frontier models approaching industry-expert quality on well-specified deliverables.
[Pearce and Song 2024] Pearce, Song. Reconciling kaplan and chinchilla scaling laws. Transactions on Machine Learning Research (TMLR). https://arxiv.org/abs/2406.12907
Attributes the Kaplan-Chinchilla disagreement largely to Kaplan counting non-embedding rather than total parameters at small scale; once embeddings are included and scale grows, the optimal exponent converges to the Chinchilla estimate.
[Peebles and Xie 2023] Peebles, Xie. Scalable diffusion models with transformers. https://arxiv.org/abs/2212.09748
DiT replaces the U-Net backbone in latent diffusion models with a transformer, showing that FID improves consistently as model Gflops scale, achieving state-of-the-art FID 2.27 on ImageNet 256x256.
[Penedo et al. 2023] Penedo, Malartic, Hesslow, Cojocaru, Cappelli, Alobeidli, Pannier, Almazrouei, Launay. The RefinedWeb dataset for falcon LLM: Outperforming curated corpora with web data, and web data only. https://arxiv.org/abs/2306.01116
RefinedWeb shows that aggressively filtered and deduplicated CommonCrawl web data alone, yielding five trillion tokens, can train LLMs that outperform models trained on curated corpora like The Pile.
[Penedo et al. 2024] Penedo, Kydlíček, Ben Allal, Lozhkov, Mitchell, Raffel, Von Werra, Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. https://arxiv.org/abs/2406.17557
FineWeb is a 15-trillion token pretraining dataset from 96 Common Crawl snapshots, with ablation-guided filtering and per-snapshot MinHash deduplication, that outperforms other public pretraining datasets; FineWeb-Edu is a 1.3-trillion token educational subset with strong MMLU and ARC results.
[Peng et al. 2023] Peng, Quesnelle, Fan, Shippole. YaRN: Efficient context window extension of large language models. arXiv preprint arXiv:2309.00071. https://arxiv.org/abs/2309.00071
YaRN extends RoPE-based LLaMA context windows using far fewer tokens and training steps than previous approaches, demonstrating extrapolation beyond the fine-tuning length.
[Peng et al. 2024] Peng, Michael, Sleight, Perez, Sharma. Rapid response: Mitigating LLM jailbreaks with a few examples. arXiv preprint arXiv:2411.07494. https://arxiv.org/abs/2411.07494
RapidResponseBench benchmarks defenses that fine-tune an LLM input classifier on proliferated jailbreak examples, reducing attack success rate by over 240x after seeing just one jailbreak per strategy.
[Perez et al. 2022] Perez, Huang, Song, Cai, Ring, Aslanides, Glaese, McAleese, Irving. Red teaming language models with language models. arXiv preprint arXiv:2202.03286. https://arxiv.org/abs/2202.03286
This paper introduces LM-based red-teaming, using one language model to automatically generate test cases that elicit harmful outputs from a target LM, uncovering tens of thousands of failures in a 280B-parameter chatbot.
[Petrov et al. 2023] Petrov, La Malfa, Torr, Bibi. Language model tokenizers introduce unfairness between languages. https://proceedings.neurips.cc/paper_files/paper/2023/hash/74bb24dca8334adce292883b4b651eda-Abstract-Conference.html
Tokenizers systematically encode the same text into up to 15 times more tokens for non-English languages, creating unfair cost, latency, and context-length disparities across language communities.
[pgvector 2025] pgvector. pgvector: Open-source vector similarity search for postgres. https://github.com/pgvector/pgvector
pgvector is an open-source PostgreSQL extension that adds vector similarity search, enabling nearest-neighbor retrieval directly in Postgres.
[Phan et al. 2025] Phan, Gatti, Han, Li, Hu, others. Humanity's last exam. arXiv preprint arXiv:2501.14249. https://arxiv.org/abs/2501.14249
Humanity's Last Exam is a broad expert-level multimodal benchmark of closed-ended academic questions designed to remain hard after MMLU-style tests saturate.
[Physical Intelligence 2025] Physical Intelligence. pi-0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054. https://arxiv.org/abs/2504.16054
pi-0.5 is a VLA model that uses co-training on heterogeneous data sources (other robots, web data, semantic subtasks, verbal instructions) to enable open-world generalization for long-horizon household manipulation tasks in unseen homes.
[Pinecone 2025] Pinecone. Pinecone: Managed vector database for AI. https://www.pinecone.io/
Pinecone is a managed vector database service providing low-latency similarity search and metadata filtering for building RAG pipelines and AI applications at scale.
[Podell et al. 2024] Podell, English, Lacey, Blattmann, Dockhorn, Müller, Penna, Rombach. SDXL: Improving latent diffusion models for high-resolution image synthesis. https://arxiv.org/abs/2307.01952
SDXL is a latent diffusion model for text-to-image synthesis featuring a 2.6B-parameter UNet with dual text encoders, novel size and crop conditioning, multi-aspect training, and an optional refinement stage.
[Polu and Sutskever 2020] Polu, Sutskever. Generative language modeling for automated theorem proving. arXiv preprint arXiv:2009.03393. https://arxiv.org/abs/2009.03393
GPT-f applies transformer language models to Metamath proof search and contributes short proofs accepted into the formal mathematics library.
[Poly 2026] Poly. Durable execution: The key to harnessing AI agents in production. https://www.inngest.com/blog/durable-execution-key-to-harnessing-ai-agents
Inngest blog post arguing that durable execution, a programming model guaranteeing code completion despite failures, is the key infrastructure primitive for running AI agents reliably in production.
[Polyak and others 2024] Polyak, others. Movie gen: a cast of media foundation models. arXiv preprint arXiv:2410.13720. https://arxiv.org/abs/2410.13720
Meta's media foundation models, whose video model carries a 73,000-token context for sixteen seconds of high-definition footage, against roughly four thousand latent tokens for one megapixel image.
[Pope et al. 2022] Pope, Douglas, Chowdhery, Devlin, Bradbury, Levskaya, Heek, Xiao, Agrawal, Dean. Efficiently scaling transformer inference. arXiv preprint arXiv:2211.05102. https://arxiv.org/abs/2211.05102
This paper presents an analytical partitioning framework and low-level optimizations for efficient Transformer inference on TPU v4 slices, achieving 29ms per token and 76
[Press and Wolf 2017] Press, Wolf. Using the output embedding to improve language models. https://arxiv.org/abs/1608.05859
Tying the input and output embedding matrices in neural language models reduces perplexity and can cut translation model parameter count to less than half with no performance loss.
[Press et al. 2022] Press, Smith, Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. https://arxiv.org/abs/2108.12409
ALiBi replaces positional embeddings with per-head linear distance penalties on attention scores, enabling transformers trained on short sequences to extrapolate to longer ones at inference time with no added runtime cost.
[Python Software Foundation 2024] Python Software Foundation. pickle: Python object serialization (module documentation). https://docs.python.org/3/library/pickle.html
The pickle module documentation, whose top-of-page warning states that unpickling untrusted data can execute arbitrary code by design, the mechanism behind malicious checkpoints.
[PyTorch Foundation 2024] PyTorch Foundation. ExecuTorch. https://github.com/pytorch/executorch
ExecuTorch is PyTorch's official framework for deploying AI models on-device across mobile, embedded, and edge hardware.
[Qian et al. 2021] Qian, Zhou, Bao, Wang, Qiu, Zhang, Yu, Li. Glancing transformer for non-autoregressive neural machine translation. https://arxiv.org/abs/2008.07905
GLAT introduces a Glancing Language Model training strategy that enables non-autoregressive neural machine translation in a single parallel decoding pass, closing the quality gap to autoregressive Transformer to within 0.25-0.9 BLEU at 8x-15x speedup.
[Qian et al. 2025] Qian, Acikgoz, He, Wang, Chen, Hakkani-Tür, Tur, Ji. ToolRL: Reward is all tool learning needs. arXiv preprint arXiv:2504.13958. https://arxiv.org/abs/2504.13958
ToolRL presents the first systematic study of reward design for tool-integrated reasoning in LLMs, using GRPO with fine-grained, decomposed rewards to outperform SFT by 15
[Qin et al. 2023] Qin, Liang, Ye, Zhu, Yan, Lu, Lin, Cong, Tang, Qian, others. ToolLLM: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789. https://arxiv.org/abs/2307.16789
ToolLLM builds ToolBench from 16,464 real-world APIs and introduces ToolEval to evaluate whether language models can plan and execute API-use tasks.
[Qin et al. 2024] Qin, Li, He, Zhang, Wu, Zheng, Xu. Mooncake: a kvcache-centric disaggregated architecture for LLM serving. arXiv preprint arXiv:2407.00079. https://arxiv.org/abs/2407.00079
Mooncake is a KVCache-centric disaggregated LLM serving platform for Kimi that separates prefill and decoding clusters and uses idle CPU/DRAM/SSD to cache KVCache, achieving up to 525
[Qin et al. 2025] Qin, Ye, Fang, Wang, Liang, Tian, Zhang, Li, Li, Huang, Zhong, Li, Yang, Miao, Lin, Liu, Jiang, Ma, Li, Xiao, Cai, Li, Zheng, Jin, Li, Zhou, Wang, Chen, Li, Yang, Liu, Lin, Peng, Liu, Shi. UI-TARS: Pioneering automated GUI interaction with native agents. arXiv preprint arXiv:2501.12326. https://arxiv.org/abs/2501.12326
A native GUI agent trained end to end, screenshots in and human-like actions out with no commercial-model wrapper, refined by reinforcement learning through hundreds of live virtual machines.
[Qu et al. 2021] Qu, Ding, Liu, Liu, Ren, Zhao, Dong, Wu, Wang. RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering. https://arxiv.org/abs/2010.08191
RocketQA improves dual-encoder dense passage retrieval for open-domain QA via three training strategies: cross-batch negatives, denoised hard negatives, and cross-encoder-based data augmentation.
[Qualcomm Technologies, Inc. 2024] Qualcomm Technologies, Inc.. Qualcomm AI engine direct SDK. https://www.qualcomm.com/developer/software/qualcomm-ai-engine-direct-sdk
Qualcomm AI Engine Direct SDK provides a unified API and hardware-specific libraries for deploying full-stack AI inference on Snapdragon processors.
[Qwen Team 2025] Qwen Team. Qwen3-next-80B-A3B-instruct model card. https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
Qwen3-Next is an 80B-total, 3B-active MoE model whose 48 layers alternate three Gated DeltaNet linear-attention layers with one gated full-attention layer.
[Rabe and Staats 2021] Rabe, Staats. Self-attention does not need O(n^2) memory. arXiv preprint arXiv:2112.05682. https://arxiv.org/abs/2112.05682
[Rack2Cloud 2026] Rack2Cloud. Kubernetes day 2 failures: 5 incidents and the metrics that predict them. https://www.rack2cloud.com/kubernetes-day-2-failures/
A practitioner guide cataloguing five common Kubernetes Day 2 failure modes and identifying one predictive metric for each to enable early detection.
[Radford et al. 2021] Radford, Kim, Hallacy, Ramesh, Goh, Agarwal, Sastry, Askell, Mishkin, Clark, others. Learning transferable visual models from natural language supervision. https://arxiv.org/abs/2103.00020
CLIP trains an image encoder and text encoder jointly on 400 million image-text pairs via contrastive learning, enabling zero-shot transfer to downstream visual tasks through natural language.
[Radford et al. 2023] Radford, Kim, Xu, Brockman, McLeavey, Sutskever. Robust speech recognition via large-scale weak supervision. https://arxiv.org/abs/2212.04356
Whisper trains an encoder-decoder Transformer on 680,000 hours of weakly supervised multilingual audio and achieves near-human robustness on ASR in a zero-shot setting without dataset-specific fine-tuning.
[Rafailov et al. 2023] Rafailov, Sharma, Mitchell, Ermon, Manning, Finn. Direct preference optimization: Your language model is secretly a reward model. https://arxiv.org/abs/2305.18290
DPO replaces RLHF's explicit reward model and RL loop with a simple binary cross-entropy loss, extracting the optimal policy directly from human preference data.
[Raffel et al. 2020] Raffel, Shazeer, Roberts, Lee, Narang, Matena, Zhou, Li, Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research. https://www.jmlr.org/papers/v21/20-074.html
T5 introduces a unified text-to-text framework that casts all NLP tasks into the same format and systematically studies pre-training objectives, architectures, and data scale to achieve state-of-the-art results.
[Ragan-Kelley et al. 2013] Ragan-Kelley, Barnes, Adams, Paris, Durand, Amarasinghe. Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. https://dl.acm.org/doi/10.1145/2491956.2462176
Halide separates what to compute from how to schedule it onto the machine, the founding idea behind every tensor compiler's search over schedules.
[Rajbhandari et al. 2019] Rajbhandari, Rasley, Ruwase, He. ZeRO: Memory optimizations toward training trillion parameter models. https://arxiv.org/abs/1910.02054
ZeRO partitions optimizer states, gradients, and parameters across data-parallel workers to eliminate memory redundancy, enabling training of models beyond 100B parameters with super-linear throughput scaling.
[Rajbhandari et al. 2021] Rajbhandari, Ruwase, Rasley, Smith, He. ZeRO-infinity: Breaking the GPU memory wall for extreme scale deep learning. https://arxiv.org/abs/2104.07857
ZeRO-Infinity is a heterogeneous training system that offloads model states to CPU and NVMe memory, enabling training of models with tens of trillions of parameters on existing GPU clusters without model code refactoring.
[Rajput et al. 2023] Rajput, Mehta, Singh, Keshavan, Vu, Heldt, Hong, Tay, Tran, Samost, Kula, Chi, Sathiamoorthy. Recommender systems with generative retrieval. https://arxiv.org/abs/2305.05065
TIGER quantizes each item's content embedding into a semantic ID and trains a sequence model to generate the next item's ID directly, folding the retrieval index into the model's weights.
[Ramesh et al. 2022] Ramesh, Dhariwal, Nichol, Chu, Chen. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125. https://arxiv.org/abs/2204.06125
DALL-E 2 (unCLIP) generates images from text by chaining a prior that maps text to CLIP image embeddings with a diffusion decoder that inverts those embeddings into 1024x1024 images.
[Rand et al. 2025] Rand, Seel, Wiser. Queued up: 2025 edition. Characteristics of power plants seeking transmission interconnection. https://emp.lbl.gov/publications/queued-2025-edition-characteristics
[Rao and Ballard 1999] Rao, Ballard. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects. Nature Neuroscience.
Rao and Ballard propose predictive coding, where higher cortical areas predict lower-level activity and only the prediction errors propagate forward.
[Raschka 2024] Raschka. Build a large language model (from scratch). Manning Publications. https://www.manning.com/books/build-a-large-language-model-from-scratch
A hands-on book guiding readers to implement LLM attention mechanisms and GPT-style transformer architectures from scratch, covering training, fine-tuning, and instruction following.
[Rasmussen et al. 2025] Rasmussen, Paliychuk, Beauvais, Ryan, Chalef. Zep: a temporal knowledge graph architecture for agent memory. arXiv preprint arXiv:2501.13956. https://arxiv.org/abs/2501.13956
Memory as a temporal knowledge graph: facts carry validity intervals and are invalidated rather than overwritten, so the store can answer what was true when.
[Ratner et al. 2016] Ratner, De Sa, Wu, Selsam, Ré. Data programming: Creating large training sets, quickly. Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/1605.07723
Data programming lets users write noisy labeling functions instead of hand-labeling examples; a generative model denoises their outputs to produce large training sets programmatically.
[Rebedea et al. 2023] Rebedea, Dinu, Sreedhar, Parisien, Cohen. NeMo guardrails: a toolkit for controllable and safe LLM applications with programmable rails. arXiv preprint arXiv:2310.10501. https://arxiv.org/abs/2310.10501
NeMo Guardrails is an open-source toolkit that adds programmable, runtime-defined guardrails to LLM applications using Colang, a custom dialogue-flow language, without modifying the underlying model.
[Rehberger 2024] Rehberger. SpAIware: Spyware injection into ChatGPT's long-term memory. https://embracethered.com/blog/posts/2024/chatgpt-macos-app-persistent-data-exfiltration/
A prompt injection that survives the session: exfiltration instructions planted in assistant memory leaked every subsequent conversation until the memory was found and deleted.
[Reimers and Gurevych 2019] Reimers, Gurevych. Sentence-BERT: Sentence embeddings using siamese BERT-networks. https://arxiv.org/abs/1908.10084
SBERT fine-tunes BERT with siamese and triplet networks to produce semantically meaningful sentence embeddings comparable via cosine-similarity, reducing pairwise similarity search from 65 hours to 5 seconds.
[Rein et al. 2023] Rein, Hou, Stickland, Petty, Pang, Dirani, Michael, Bowman. GPQA: a graduate-level google-proof Q&a benchmark. arXiv preprint arXiv:2311.12022. https://arxiv.org/abs/2311.12022
GPQA is a 448-question graduate-level benchmark in biology, physics, and chemistry where PhD experts reach 65
[Ren et al. 2025] Ren, Shao, Song, Xin, Wang, Zhao, Zhang, Fu, Zhu, Yang, Wu, Gou, Ma, Tang, Liu, Gao, Guo, Ruan. DeepSeek-prover-V2: Advancing formal mathematical reasoning via reinforcement learning for subgoal decomposition. arXiv preprint arXiv:2504.21801. https://arxiv.org/abs/2504.21801
DeepSeek-Prover-V2 combines informal and formal reasoning for Lean 4 theorem proving, using recursive decomposition and RL to reach strong MiniF2F and PutnamBench results.
[Replit 2025] Replit. Checkpoints and rollbacks. https://docs.replit.com/replitai/checkpoints-and-rollbacks
Replit's Checkpoints and Rollbacks docs page describes how the Replit Agent automatically saves project snapshots at development milestones, enabling one-click restore of full project state and AI conversation context.
[Replit 2025] Replit. Inside replit's snapshot engine. https://blog.replit.com/inside-replits-snapshot-engine
[Replit 2025] Replit. Finding and solving memory leaks. https://blog.replit.com/finding-and-solving-memory-leaks
[ReversingLabs 2025] ReversingLabs. nullifAI: malicious ML models evade Hugging Face picklescan. https://www.reversinglabs.com/blog/rl-identifies-malware-ml-model-hosted-on-hugging-face
Malicious models that evaded the hub's pickle scanner by using an unparsed compression format and placing the payload before a deliberately broken instruction, showing scanning is not a guarantee.
[Riedl and Desai 2025] Riedl, Desai. AI agents and the law. arXiv preprint arXiv:2508.08544. https://arxiv.org/abs/2508.08544
Compares computer science's and law's notions of an agent, arguing the technical framing undertheorizes loyalty and third-party obligations, exactly what agency law exists to govern.
[Rissanen 1978] Rissanen. Modeling by shortest data description. Automatica.
Rissanen introduces the Minimum Description Length principle, framing model selection as choosing the model that most compresses the data.
[Rombach et al. 2022] Rombach, Blattmann, Lorenz, Esser, Ommer. High-resolution image synthesis with latent diffusion models. https://arxiv.org/abs/2112.10752
Latent Diffusion Models (LDMs) run diffusion in a pretrained autoencoder's latent space, cutting training and inference cost while adding cross-attention conditioning for text-to-image and other tasks.
[Rosenblatt 1958] Rosenblatt. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological Review.
Rosenblatt introduces the perceptron, a trainable linear classifier with a weight-update learning rule, an early foundation of neural networks.
[Rosenfeld et al. 2020] Rosenfeld, Rosenfeld, Belinkov, Shavit. A constructive prediction of the generalization error across scales. https://arxiv.org/abs/1909.12673
Proposes the parametric envelope form of the generalization error as a power law in model size and data size plus an irreducible floor, fit across vision and language, predating its language-model specialization.
[Röttger et al. 2024] Röttger, Kirk, Vidgen, Attanasio, Bianchi, Hovy. XSTest: a test suite for identifying exaggerated safety behaviours in large language models. https://arxiv.org/abs/2308.01263
XSTEST is a 250-prompt test suite that identifies exaggerated safety behaviors in LLMs, where models refuse safe prompts due to lexical overlap with unsafe ones.
[Rouhani and others 2023] Rouhani, others. Microscaling data formats for deep learning. https://arxiv.org/abs/2310.10537
The OCP Microscaling (MX) proposal from AMD, Arm, Intel, Meta, Microsoft, NVIDIA, and Qualcomm pairs narrow floating-point and integer element types with a shared per-block scale, and shows MX formats, including MXFP4, working for inference and training with minimal accuracy loss.
[Rumelhart et al. 1986] Rumelhart, Hinton, Williams. Learning representations by back-propagating errors. Nature. https://doi.org/10.1038/323533a0
[Russinovich et al. 2024] Russinovich, Salem, Eldan. Great, now write an article about that: The crescendo multi-turn LLM jailbreak attack. arXiv preprint arXiv:2404.01833. https://arxiv.org/abs/2404.01833
Crescendo is a multi-turn LLM jailbreak that escalates from benign prompts using the model's own prior outputs, bypassing safety alignment on GPT-4, Gemini-Pro, and other models with high attack success rates.
[Saharia et al. 2022] Saharia, Chan, Saxena, others. Photorealistic text-to-image diffusion models with deep language understanding. https://arxiv.org/abs/2205.11487
Imagen combines a frozen T5-XXL text encoder with a cascade of diffusion models, finding that scaling the language model improves text-to-image quality more than scaling the image model, achieving FID 7.27 on COCO.
[Sahoo et al. 2024] Sahoo, Arriola, Schiff, Gokaslan, Marroquin, Chiu, Rush, Kuleshov. Simple and effective masked diffusion language models. https://arxiv.org/abs/2406.07524
MDLM shows that masked discrete diffusion with a Rao-Blackwellized objective and modern training recipes closes most of the perplexity gap between diffusion and autoregressive language models.
[Sakana AI 2025] Sakana AI. The AI CUDA engineer: Agentic CUDA kernel discovery, optimization and composition. https://sakana.ai/ai-cuda-engineer/
[Salesforce 2025] Salesforce. 2025 holiday shopping data. https://www.salesforce.com/news/stories/2025-holiday-shopping-data/
[Salimans et al. 2017] Salimans, Ho, Chen, Sidor, Sutskever. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864. https://arxiv.org/abs/1703.03864
Evolution Strategies (ES) scaled to 1,440 parallel workers via shared random seeds achieves competitive results with policy-gradient RL on MuJoCo and Atari without backpropagation or value functions.
[Salimans and Ho 2022] Salimans, Ho. Progressive distillation for fast sampling of diffusion models. https://arxiv.org/abs/2202.00512
Repeatedly distills a diffusion sampler into a student that needs half the steps, taking generation from thousands of steps down to as few as four at no more than the original training cost.
[Sambasivan et al. 2021] Sambasivan, Kapania, Highfill, Akrong, Paritosh, Aroyo. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. https://doi.org/10.1145/3411764.3445518
Sambasivan et al. document data cascades in high-stakes AI, where upstream data problems compound into large downstream model failures.
[San Roman et al. 2024] San Roman, Fernandez, Défossez, Furon, Tran, Elsahar. Proactive detection of voice cloning with localized watermarking. https://arxiv.org/abs/2401.17264
AudioSeal is the first audio watermarking system designed for localized, sample-level detection of AI-generated speech, using a jointly trained generator/detector architecture that achieves up to two orders of magnitude faster detection than prior methods.
[Santhanam et al. 2022] Santhanam, Khattab, Saad-Falcon, Potts, Zaharia. ColBERTv2: Effective and efficient retrieval via lightweight late interaction. https://arxiv.org/abs/2112.01488
ColBERTv2 combines residual compression and denoised supervision to cut late-interaction retrieval storage by 6-10x while achieving state-of-the-art quality within and outside its training domain.
[Sardana et al. 2024] Sardana, Portes, Doubov, Frankle. Beyond chinchilla-optimal: Accounting for inference in language model scaling laws. https://arxiv.org/abs/2401.00448
This paper extends Chinchilla scaling laws to account for inference costs, showing that practitioners expecting large inference demand should train smaller models on more tokens than Chinchilla-optimal.
[Sardana et al. 2024] Sardana, Portes, Doubov, Frankle. Beyond chinchilla-optimal: Accounting for inference in language model scaling laws. https://arxiv.org/abs/2401.00448
This paper extends Chinchilla scaling laws to account for inference costs, showing that developers expecting large inference demand ( 1B requests) should train smaller models on more tokens than Chinchilla-optimal.
[Sauer et al. 2024] Sauer, Lorenz, Blattmann, Rombach. Adversarial diffusion distillation. https://arxiv.org/abs/2311.17042
Combines score distillation with an adversarial loss for one-to-four-step sampling, the method behind SDXL-Turbo, reaching real-time synthesis that matches its teacher within four steps.
[Savinov et al. 2022] Savinov, Chung, Binkowski, Elsen, Oord. Step-unrolled denoising autoencoders for text generation. https://arxiv.org/abs/2112.06749
SUNDAE proposes a non-autoregressive text generative model that iteratively denoises a sequence of tokens using unrolled denoising during training, achieving state-of-the-art results among non-autoregressive methods on WMT'14 translation.
[Saxena 2023] Saxena. Prompt lookup decoding. https://github.com/apoorvumang/prompt-lookup-decoding
Prompt lookup decoding replaces the draft model in speculative decoding with simple n-gram matching against the prompt, yielding 2-4x speedups on input-grounded tasks with no change to output quality; it ships in transformers and vLLM.
[Schaeffer et al. 2023] Schaeffer, Miranda, Koyejo. Are emergent abilities of large language models a mirage?. https://arxiv.org/abs/2304.15004
Apparent emergent abilities in LLMs are artifacts of nonlinear or discontinuous evaluation metrics, not fundamental changes in model behavior with scale.
[Schick et al. 2023] Schick, Dwivedi-Yu, Dessì, Raileanu, Lomeli, Zettlemoyer, Cancedda, Scialom. Toolformer: Language models can teach themselves to use tools. https://arxiv.org/abs/2302.04761
Toolformer trains a language model via self-supervised API-call filtering to decide when and how to invoke external tools, achieving strong zero-shot performance without human annotation.
[Schulman et al. 2017] Schulman, Wolski, Dhariwal, Radford, Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. https://arxiv.org/abs/1707.06347
PPO introduces a clipped surrogate objective for policy gradient reinforcement learning that achieves TRPO-level reliability with simpler first-order optimization and better sample complexity.
[Schulman and Thinking Machines Lab 2025] Schulman, Thinking Machines Lab. LoRA without regret. https://thinkingmachines.ai/blog/lora/
Large-scale experiments showing LoRA matches full fine-tuning when applied to all linear layers (MLP and MoE included) at about ten times the full-FT learning rate; attention-only targeting underperforms, and LoRA matches full FT for RL post-training even at rank 1.
[Schultz et al. 1997] Schultz, Dayan, Montague. A neural substrate of prediction and reward. Science.
Schultz, Dayan, and Montague show that midbrain dopamine neurons encode a reward prediction error, linking neuroscience to temporal-difference reinforcement learning.
[Schuman 2026] Schuman. What really caused that AWS outage in december?. https://www.computerworld.com/article/4136512/what-really-caused-that-aws-outage-in-december.html
Computerworld's follow-up on the December 2025 AWS Cost Explorer outage: AWS confirmed its Kiro agent deleted and recreated an environment but attributed the root cause to a misconfigured role, a framing the author disputes.
[Sculley et al. 2015] Sculley, Holt, Golovin, Davydov, Phillips, Ebner, Chaudhary, Young, Crespo, Dennison. Hidden technical debt in machine learning systems. https://papers.nips.cc/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
Sculley et al. describe hidden technical debt in ML systems, including boundary erosion, entanglement, data dependencies, and undeclared consumers.
[Seamless Communication 2023] Seamless Communication. SeamlessM4T: Massively multilingual & multimodal machine translation. arXiv preprint arXiv:2308.11596. https://arxiv.org/abs/2308.11596
SeamlessM4T is a single model supporting ASR, speech-to-speech, speech-to-text, text-to-speech, and text-to-text translation across up to 100 languages, trained on 406,000 hours of multilingual speech data.
[Seamless Communication 2023] Seamless Communication. Seamless: Multilingual expressive and streaming speech translation. arXiv preprint arXiv:2312.05187. https://arxiv.org/abs/2312.05187
Seamless introduces SeamlessM4T v2, SeamlessExpressive, and SeamlessStreaming: a family of multilingual models enabling real-time expressive speech-to-speech translation with prosody and vocal style preservation.
[SemiAnalysis 2024] SemiAnalysis. MI300X vs H100 vs H200 benchmark part 1: Training, CUDA moat still alive. https://semianalysis.com/2024/12/22/mi300x-vs-h100-vs-h200-benchmark-part-1-training/
[Sen et al. 2026] Sen, Kasturi, Lumer, Gulati, Subbiah. Is grep all you need? How agent harnesses reshape agentic search. arXiv preprint arXiv:2605.15184. https://arxiv.org/abs/2605.15184
A controlled comparison of retrieval strategies inside agent harnesses: grep-style search over live state generally yields higher accuracy than vector retrieval, with the margin depending on the harness and tool-calling style.
[Senarath and Dissanayaka 2025] Senarath, Dissanayaka. OAuth 2.0 extension: On-behalf-of user authorization for AI agents (internet-draft, expired). https://datatracker.ietf.org/doc/draft-oauth-ai-agents-on-behalf-of-user/
[Sennrich et al. 2015] Sennrich, Haddow, Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909. https://arxiv.org/abs/1508.07909
This paper proposes using BPE to segment rare and unknown words into subword units, enabling open-vocabulary neural machine translation without back-off dictionaries.
[Settles 2009] Settles. Active learning literature survey. https://minds.wisconsin.edu/handle/1793/60660
[Seung et al. 1992] Seung, Sompolinsky, Tishby. Statistical mechanics of learning from examples. Physical Review A.
Seung, Sompolinsky, and Tishby apply statistical mechanics to learning, deriving generalization error and learning curves as a function of training-set size.
[Seyedi and NVIDIA 2025] Seyedi, NVIDIA. Scaling AI factories with co-packaged optics for better power efficiency. https://developer.nvidia.com/blog/scaling-ai-factories-with-co-packaged-optics-for-better-power-efficiency/
NVIDIA blog post explaining how co-packaged optics (CPO) reduce power consumption in AI data center networking to scale LLM training clusters.
[SGLang Team 2025] SGLang Team. Towards deterministic inference in sglang and reproducible RL training. https://lmsys.org/blog/2025-09-22-sglang-deterministic/
[Shah et al. 2024] Shah, Bikshandi, Zhang, Thakkar, Ramani, Dao. FlashAttention-3: Fast and accurate attention with asynchrony and low-precision. https://arxiv.org/abs/2407.08608
FlashAttention-3 exploits Hopper GPU asynchrony via warp-specialization and FP8 low-precision to achieve 1.5-2x speedup over FlashAttention-2, reaching 740 TFLOPs/s in FP16 and 1.2 PFLOPs/s in FP8.
[Shankar et al. 2022] Shankar, Garcia, Hellerstein, Parameswaran. Operationalizing machine learning: An interview study. arXiv preprint arXiv:2209.09125. https://arxiv.org/abs/2209.09125
An interview study of 18 ML engineers identifies three MLOps success variables (Velocity, Validation, Versioning) and documents practices and pain points for deploying and sustaining ML pipelines in production.
[Shannon 1948] Shannon. A mathematical theory of communication. Bell System Technical Journal.
Shannon founds information theory, defining entropy as the measure of information and establishing the limits of compression and reliable communication over noisy channels.
[Shao et al. 2024] Shao, Wang, Zhu, Xu, Song, Bi, Zhang, Zhang, Li, Wu, Guo. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. https://arxiv.org/abs/2402.03300
DeepSeekMath combines a curated 120B-token math corpus with GRPO, a PPO variant that removes the critic and normalizes rewards within sampled groups.
[Shao et al. 2025] Shao, Li, Xin, Geng, Wang, Oh, Du, Lambert, Min, Krishna, Tsvetkov, Hajishirzi, Koh, Zettlemoyer. Spurious rewards: Rethinking training signals in RLVR. arXiv preprint arXiv:2506.10947. https://arxiv.org/abs/2506.10947
RLVR with random or incorrect rewards recovers most of the ground-truth gain on Qwen2.5-Math-7B (+21.4 vs +29.1 points on MATH-500) yet fails on Llama and OLMo, suggesting GRPO amplifies pretrained behaviors rather than teaching new reasoning.
[Sharma and Kaplan 2020] Sharma, Kaplan. A neural scaling law from the dimension of the data manifold. arXiv preprint arXiv:2004.10802. https://arxiv.org/abs/2004.10802
Derives the scaling exponent from the intrinsic dimension of the data manifold, predicting that loss falls as a power law whose rate is set by how many dimensions the data effectively occupies.
[Sharma et al. 2024] Sharma, Tong, Korbak, Duvenaud, Askell, Bowman, Cheng, Durmus, Hatfield-Dodds, Johnston, Kravec, Maxwell, McCandlish, Ndousse, Rausch, Schiefer, Yan, Zhang, Perez. Towards understanding sycophancy in language models. https://arxiv.org/abs/2310.13548
[Sharma et al. 2025] Sharma, Tong, Mu, Wei, Kruthoff, Goodfriend, Ong, Peng, Agarwal, Anil, others. Constitutional classifiers: Defending against universal jailbreaks across thousands of hours of red teaming. arXiv preprint arXiv:2501.18837. https://arxiv.org/abs/2501.18837
Constitutional Classifiers train LLM safeguard classifiers on constitution-guided synthetic data to block universal jailbreaks, achieving over 95
[Shazeer et al. 2017] Shazeer, Mirhoseini, Maziarz, Davis, Le, Hinton, Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. https://arxiv.org/abs/1701.06538
This paper introduces a Sparsely-Gated MoE layer with up to thousands of feed-forward experts and a trainable gating network, achieving over 1000x model capacity gains with minor computational overhead on language modeling and translation tasks.
[Shazeer 2019] Shazeer. Fast transformer decoding: One write-head is all you need. https://arxiv.org/abs/1911.02150
Multi-query attention (MQA) shares keys and values across all attention heads, cutting memory bandwidth for incremental decoding with only minor quality loss.
[Shazeer 2020] Shazeer. GLU variants improve transformer. https://arxiv.org/abs/2002.05202
This paper proposes GLU variants such as SwiGLU and GEGLU as replacements for the ReLU activation in the Transformer FFN sublayer, finding they improve perplexity and downstream task quality.
[Shehabi et al. 2024] Shehabi, Smith, Hubbard, Newkirk, Lei, Siddik, Holecek, Koomey, Masanet, Sartor. 2024 united states data center energy usage report. https://escholarship.org/uc/item/32d6m0d1
[Sheng et al. 2023] Sheng, Cao, Li, Hooper, Lee, Yang, Chou, Zhu, Zheng, Keutzer, Gonzalez, Stoica. S-LoRA: Serving thousands of concurrent LoRA adapters. arXiv preprint arXiv:2311.03285. https://arxiv.org/abs/2311.03285
[Sheng et al. 2025] Sheng, Zhang, Ye, Wu, Zhang, Zhang, Peng, Lin, Wu. HybridFlow: a flexible and efficient RLHF framework. https://arxiv.org/abs/2409.19256
HybridFlow combines single-controller and multi-controller paradigms for RLHF to deliver flexible dataflow representation and efficient distributed execution, with 1.53x to 20.57x throughput gains over prior systems.
[Shi et al. 2017] Shi, Karpathy, Fan, Hernandez, Liang. World of bits: An open-domain platform for web-based agents. https://proceedings.mlr.press/v70/shi17a.html
The origin of GUI agents as a research line: agents perceive pixels and DOM and act with mouse and keyboard, plus the MiniWoB task suite the field trained on for years.
[Shi et al. 2023] Shi, Ajith, Xia, Huang, Liu, Blevins, Chen, Zettlemoyer. Detecting pretraining data from large language models. arXiv preprint arXiv:2310.16789. https://arxiv.org/abs/2310.16789
MIN-K
[Shi et al. 2024] Shi, Han, Wang, Doucet, Titsias. Simplified and generalized masked diffusion for discrete data. https://arxiv.org/abs/2406.04329
MD4 simplifies masked diffusion for discrete data by showing the continuous-time ELBO reduces to a weighted integral of cross-entropy losses, improving perplexity and image modeling over prior discrete diffusion models.
[Shi et al. 2025] Shi, Ma, Liang, Diao, Ma, Vosoughi. Judging the judges: a systematic study of position bias in LLM-as-a-judge. https://aclanthology.org/2025.ijcnlp-long.18/
A systematic study of position bias in LLM-as-a-Judge, introducing three metrics across 15 judges and 150,000+ instances, finding bias correlates most strongly with solution quality gaps.
[Shilov 2026] Shilov. TSMC 'super carrier' CoWoS interposer enables 9-reticle packages with 12 HBM4 stacks. https://www.tomshardware.com/tech-industry/tsmc-super-carrier-cowos-interposer-gets-bigger-enabling-massive-ai-chips-to-reach-9-reticle-sizes-with-12-hbm4-stacks
[Shinn et al. 2023] Shinn, Cassano, Berman, Gopinath, Narasimhan, Yao. Reflexion: Language agents with verbal reinforcement learning. arXiv preprint arXiv:2303.11366. https://arxiv.org/abs/2303.11366
Reflexion reinforces LLM agents without weight updates by storing verbal self-reflections in episodic memory, achieving 91
[Shoeybi et al. 2019] Shoeybi, Patwary, Puri, LeGresley, Casper, Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism. https://arxiv.org/abs/1909.08053
Megatron-LM introduces a simple intra-layer tensor parallelism approach for PyTorch transformers, scaling GPT-2 and BERT models to 8.3 billion parameters across 512 GPUs with 76
[Shokri et al. 2017] Shokri, Stronati, Song, Shmatikov. Membership inference attacks against machine learning models. https://arxiv.org/abs/1610.05820
This paper introduces membership inference attacks that use shadow training to determine, via black-box API access, whether a record was in a model's training dataset.
[Shumailov et al. 2024] Shumailov, Shumaylov, Zhao, Papernot, Anderson, Gal. AI models collapse when trained on recursively generated data. Nature. https://www.nature.com/articles/s41586-024-07566-y
Training language models recursively on their own synthetic outputs causes progressive model collapse, compounding across generations until the output distribution narrows to noise.
[Si et al. 2025] Si, Li, Backes, Zhang. Excessive reasoning attack on reasoning llms. arXiv preprint arXiv:2506.14374. https://arxiv.org/abs/2506.14374
This attack crafts inputs that trigger excessive reasoning and delayed termination in reasoning models, increasing reasoning length by several times while preserving utility.
[Silent Data Corruption Study 2026] Silent Data Corruption Study. The anatomy of silent data corruption: GPU error pattern study and modeling guidance. arXiv preprint arXiv:2605.04213. https://arxiv.org/abs/2605.04213
A gate-level fault injection study on a production GPU characterizes SDC patterns, finding NaN/INF account for only 1.01
[Singh et al. 2025] Singh, Magazine, Pandya, Nambi. Agentic reasoning and tool integration for llms via reinforcement learning. arXiv preprint arXiv:2505.01441. https://arxiv.org/abs/2505.01441
ARTIST is a unified framework that couples agentic reasoning, reinforcement learning via GRPO, and dynamic tool integration for LLMs, achieving up to 22
[Singh 2025] Singh. AWS outage and the kiro AI bot: a post-mortem. https://singhajit.com/aws-outage-kiro-ai-bot/
A blog post analyzing a December 2025 incident where an AI agent named Kiro autonomously deleted and recreated an AWS production environment, causing a 13-hour outage due to access control failures.
[Singh et al. 2025] Singh, Ehtesham, Kumar, Khoei. Agentic retrieval-augmented generation: a survey on agentic RAG. arXiv preprint arXiv:2501.09136. https://arxiv.org/abs/2501.09136
This survey introduces a principled taxonomy of Agentic RAG architectures, tracing the evolution from naive RAG to autonomous-agent-embedded pipelines and analyzing design trade-offs, applications, and open challenges.
[Singh et al. 2025] Singh, Nan, Wang, others. The leaderboard illusion. arXiv preprint arXiv:2504.20879. https://arxiv.org/abs/2504.20879
Documents systematic distortions in Chatbot Arena rankings: undisclosed private testing of many variants with selective score retraction, asymmetric data access favoring large providers, and silent model deprecation, each of which biases the Bradley-Terry fit.
[SK hynix 2025] SK hynix. SK hynix completes world's first HBM4 development and readies mass production. https://www.prnewswire.com/news-releases/sk-hynix-completes-worlds-first-hbm4-development-and-readies-mass-production-302554538.html
SK hynix announced completion of the world's first HBM4 development and readiness for mass production of this next-generation high-bandwidth memory.
[Snap Research 2024] Snap Research. LOCOMO: Long-term conversational memory benchmark. https://snap-research.github.io/locomo/
LoCoMo is a benchmark for evaluating very long-term conversational memory of LLM agents via question answering, event summarization, and multimodal dialog generation tasks.
[Snell et al. 2024] Snell, Lee, Xu, Kumar. Scaling LLM test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314. https://arxiv.org/abs/2408.03314
This paper shows that adaptively allocating test-time compute per prompt difficulty can outperform a 14x larger pretrained model, achieving over 4x efficiency gains versus best-of-N sampling.
[Sohl-Dickstein et al. 2015] Sohl-Dickstein, Weiss, Maheswaranathan, Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. https://arxiv.org/abs/1503.03585
This paper introduces diffusion probabilistic models, which learn to reverse an iterative forward diffusion process that gradually destroys data structure, yielding a flexible and tractable generative model.
[Song and Ermon 2019] Song, Ermon. Generative modeling by estimating gradients of the data distribution. https://arxiv.org/abs/1907.05600
Song and Ermon propose Noise Conditional Score Networks (NCSN), a generative model that estimates data distribution gradients at multiple noise levels and samples via annealed Langevin dynamics, achieving state-of-the-art inception score 8.87 on CIFAR-10.
[Song et al. 2021] Song, Sohl-Dickstein, Kingma, Kumar, Ermon, Poole. Score-based generative modeling through stochastic differential equations. https://arxiv.org/abs/2011.13456
This paper unifies score-based generative models and DDPM under a continuous-time SDE framework, enabling exact likelihood computation, controllable generation, and state-of-the-art image synthesis on CIFAR-10.
[Song et al. 2021] Song, Meng, Ermon. Denoising diffusion implicit models. https://arxiv.org/abs/2010.02502
DDIM replaces the Markovian forward process of DDPM with a non-Markovian one, enabling 10x to 50x faster sampling without retraining.
[Song et al. 2023] Song, Dhariwal, Chen, Sutskever. Consistency models. https://arxiv.org/abs/2303.01469
Models that map any point on the denoising trajectory directly to its origin, enabling one-step generation with optional multi-step refinement, trainable by distillation or from scratch.
[Song et al. 2025] Song, Zhang, Luo, Gao, Xia, Luo, Li, others. Seed diffusion: a large-scale diffusion language model with high-speed inference. arXiv preprint arXiv:2508.02193. https://arxiv.org/abs/2508.02193
[Song et al. 2025] Song, Dai, Prabhu, Zhang, Shi, Li, Li, Savarese, Chen, Zhao, Xu, Xiong. CoAct-1: Computer-using multi-agent system with coding actions. arXiv preprint arXiv:2508.03923. https://arxiv.org/abs/2508.03923
CoAct-1 is a computer-use multi-agent system whose orchestrator delegates subtasks to either a GUI operator or a programmer agent, reaching 60.76
[Souly et al. 2025] Souly, Rando, others. Poisoning attacks on llms require a near-constant number of poison samples. arXiv preprint arXiv:2510.07192. https://arxiv.org/abs/2510.07192
About 250 poisoned documents backdoor models from 600M to 13B parameters regardless of clean-data volume, implying poisoning becomes relatively easier as models scale.
[Speelpenning 1980] Speelpenning. Compiling fast partial derivatives of functions given by algorithms. https://archive.org/details/compilingfastpar1002spee
[Staab et al. 2024] Staab, Vero, Balunović, Vechev. Beyond memorization: Violating privacy via inference with large language models. https://arxiv.org/abs/2310.07298
LLMs infer location, income, and demographics from ordinary text at high accuracy for a fraction of a human profiler's cost; a memory store persists exactly that inference, run continuously.
[Stanovich and West 2000] Stanovich, West. Individual differences in reasoning: Implications for the rationality debate?. Behavioral and Brain Sciences.
Stanovich and West introduce the System 1 / System 2 terminology for dual-process reasoning and analyze individual differences in human rationality.
[Starace et al. 2025] Starace, Jaffe, Sherburn, Aung, Chan, Maksin, Dias, Mays, Kinsella, Thompson, others. PaperBench: Evaluating ai's ability to replicate AI research. arXiv preprint arXiv:2504.01848. https://arxiv.org/abs/2504.01848
PaperBench evaluates AI agents on replicating AI research papers, decomposing each replication into thousands of rubric-graded subtasks and separately benchmarking the judge.
[Stray et al. 2021] Stray, Vendrov, Nixon, Adler, Hadfield-Menell. What are you optimizing for? Aligning recommender systems with human values. arXiv preprint arXiv:2107.10939. https://arxiv.org/abs/2107.10939
Frames engagement-optimized recommenders as an alignment problem deployed at planetary scale, and surveys concrete interventions for pointing them at human values instead of click proxies.
[Stripe 2025] Stripe. Developing an open standard for agentic commerce. https://stripe.com/blog/developing-an-open-standard-for-agentic-commerce
The Agentic Commerce Protocol behind Instant Checkout in ChatGPT: a shared payment token scoped to one merchant and one cart total, so the assistant never holds raw card credentials.
[Study 2025] Study. The illusion of diminishing returns: Measuring long horizon execution in llms. arXiv preprint arXiv:2509.09677. https://arxiv.org/abs/2509.09677
This paper argues that diminishing returns on short-task benchmarks mask exponential gains in long-horizon execution length, and identifies a self-conditioning failure mode where LLMs degrade when their context contains prior errors.
[Study 2026] Study. Beyond pass@1: a reliability science framework for long-horizon LLM agents. arXiv preprint arXiv:2603.29231. https://arxiv.org/abs/2603.29231
This paper proposes a reliability science framework for long-horizon LLM agents with four metrics (RDC, VAF, GDS, MOP), showing that capability and reliability rankings diverge substantially as task duration increases.
[Su et al. 2021] Su, Lu, Pan, Murtadha, Wen, Liu. RoFormer: Enhanced transformer with rotary position embedding. https://arxiv.org/abs/2104.09864
RoFormer introduces RoPE, which encodes token positions as rotation matrices in self-attention, giving sequence-length flexibility and decaying inter-token dependency with distance.
[Su et al. 2021] Su, Cao, Liu, Ou. Whitening sentence representations for better semantics and faster retrieval. arXiv preprint arXiv:2103.15316. https://arxiv.org/abs/2103.15316
This paper shows that applying a classical whitening transformation to BERT sentence embeddings corrects their anisotropy, improves semantic similarity performance, and reduces vector dimensionality for faster retrieval.
[Su and others 2022] Su, others. One embedder, any task: Instruction-finetuned text embeddings. https://arxiv.org/abs/2212.09741
INSTRUCTOR is a single text embedding model that takes task instructions at inference time to generate task- and domain-aware embeddings without further fine-tuning, achieving state-of-the-art results across 70 diverse tasks.
[Sun 2025] Sun. Why did M2 end up as a full attention model?. https://www.minimax.io/news/why-did-m2-end-up-as-a-full-attention-model
MiniMax's pre-training lead explains why M2 dropped the hybrid lightning-attention design: hybrid deficits surfaced only at scale on multi-hop reasoning, and the inference and evaluation stack around efficient attention is not yet production-mature.
[Sutton 1988] Sutton. Learning to predict by the methods of temporal differences. Machine Learning.
Sutton introduces temporal-difference learning, which updates predictions from the difference between successive estimates rather than waiting for the final outcome.
[Swinhoe 2026] Swinhoe. TSMC announces 2026 capex spend of $56bn as CEO dismisses bubble concerns. https://www.datacenterdynamics.com/en/news/tsmc-announces-2026-capex-spend-of-56bn-after-posting-eighth-consecutive-quarter-of-growth/
[Szabo 1999] Szabo. Micropayments and mental transaction costs. https://nakamotoinstitute.org/library/micropayments-and-mental-transaction-costs/
The argument that killed 1990s micropayments: the binding cost of a tiny purchase is the decision, not the fee, and Szabo's own skepticism about delegating those decisions to agents reads as prophecy.
[Tan et al. 2024] Tan, Zeng, Tian, Liu, Yin, Jiang. Democratizing large language models via personalized parameter-efficient fine-tuning. arXiv preprint arXiv:2402.04401. https://arxiv.org/abs/2402.04401
[Tang et al. 2024] Tang, Zhao, Zhu, Xiao, Kasikci, Han. Quest: Query-aware sparsity for efficient long-context LLM inference. https://arxiv.org/abs/2406.10774
Quest keeps the full KV cache but summarizes each page by its per-channel key extremes and loads only the top-K pages critical to the current query, giving up to 2.23x self-attention speedup with negligible accuracy loss on long-dependency tasks.
[Templeton et al. 2024] Templeton, Conerly, Marcus, Lindsey, Bricken, Chen, Pearce, Citro, Ameisen, Jones, Cunningham, Turner, McDougall, MacDiarmid, Freeman, Sumers, Rees, Batson, Jermyn, Carter, Olah, Henighan. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
[Thakur et al. 2021] Thakur, Reimers, Rücklé, Srivastava, Gurevych. BEIR: a heterogenous benchmark for zero-shot evaluation of information retrieval models. https://arxiv.org/abs/2104.08663
BEIR is a heterogeneous benchmark of 18 datasets across 9 retrieval task types for zero-shot evaluation of information retrieval models, finding BM25 and re-ranking methods most robust out-of-distribution.
[The Authors Guild 2025] The Authors Guild. Bartz v. anthropic settlement: What authors need to know. https://authorsguild.org/advocacy/artificial-intelligence/what-authors-need-to-know-about-the-anthropic-settlement/
[The Hacker News 2025] The Hacker News. Experts find AI browsers can be tricked by PromptFix exploit. https://thehackernews.com/2025/08/experts-find-ai-browsers-can-be-tricked.html
Coverage of Guardio Labs' Scamlexity research: agentic browsers bought from fake storefronts and followed instructions hidden in counterfeit CAPTCHAs, spending stored payment details autonomously.
[The Linux Foundation 2025] The Linux Foundation. Linux foundation launches the Agent2Agent protocol project to enable secure, intelligent communication between AI agents. https://www.linuxfoundation.org/press/linux-foundation-launches-the-agent2agent-protocol-project-to-enable-secure-intelligent-communication-between-ai-agents
Google donates the A2A protocol to the Linux Foundation, placing agent-to-agent interoperability under vendor-neutral governance with support from more than 100 technology companies.
[The Linux Foundation 2026] The Linux Foundation. A2A protocol surpasses 150 organizations, lands in major cloud platforms, and sees enterprise production use in first year. https://www.prnewswire.com/news-releases/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year-302737641.html
A2A reached its v1.0 stable specification in April 2026 with more than 150 supporting organizations and platform integrations across Azure AI Foundry, Amazon Bedrock AgentCore, and Google Cloud.
[Thorne et al. 2018] Thorne, Vlachos, Christodoulopoulos, Mittal. FEVER: a large-scale dataset for fact extraction and verification. arXiv preprint arXiv:1803.05355. https://arxiv.org/abs/1803.05355
FEVER introduced a large claim-verification dataset where systems must retrieve evidence and label claims as supported, refuted, or not enough information.
[Tian et al. 2024] Tian, Jiang, Yuan, Peng, Wang. Visual autoregressive modeling: Scalable image generation via next-scale prediction. https://arxiv.org/abs/2404.02905
VAR redefines image autoregression as coarse-to-fine "next-scale prediction," enabling GPT-style AR models to surpass diffusion transformers in image quality, speed, and scalability on ImageNet.
[Tillet et al. 2019] Tillet, Kung, Cox. Triton: An intermediate language and compiler for tiled neural network computations. https://dl.acm.org/doi/10.1145/3315508.3329973
Triton makes the statically-shaped tile the unit of GPU programming: the programmer writes one program per tile while the compiler handles coalescing, shared memory, and intra-processor scheduling.
[Tokui et al. 2015] Tokui, Oono, Hido, Clayton. Chainer: a next-generation open source framework for deep learning. http://learningsys.org/papers/LearningSys_2015_paper_33.pdf
The paper that coined define-by-run: execute host-language code eagerly, record operations as they happen, and differentiate the recorded tape, with the case against static graphs made in print.
[Tom's Hardware 2026] Tom's Hardware. Nvidia removes rubin CPX accelerators from its roadmap. https://www.tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed
Report that NVIDIA removed the Rubin CPX prefill accelerator from its roadmap at GTC 2026, with the idea possibly returning in a later generation.
[Touvron et al. 2023] Touvron, Lavril, Izacard, Martinet, Lachaux, Lacroix, Rozière, Goyal, Hambro, Azhar, Rodriguez, Joulin, Grave, Lample. LLaMA: Open and efficient foundation language models. https://arxiv.org/abs/2302.13971
LLaMA introduces a family of open foundation language models (7B to 65B parameters) trained exclusively on public data, where smaller models trained on more tokens match or outperform larger proprietary models at inference.
[Trail of Bits 2024] Trail of Bits. A few notes on AWS nitro enclaves: Images and attestation. https://blog.trailofbits.com/2024/02/16/a-few-notes-on-aws-nitro-enclaves-images-and-attestation/
A practitioner's analysis of AWS Nitro Enclaves: what their attestation actually covers, and why isolation without memory encryption places the platform vendor inside the trust boundary.
[TrendForce 2025] TrendForce. Huawei unveils ascend 950 with in-house HBM in 2026, touts SuperPoD to rival NVIDIA. https://www.trendforce.com/news/2025/09/18/news-huawei-unveils-ascend-950-with-in-house-hbm-in-2026-touts-superpod-to-rival-nvidia/
TrendForce report on Huawei's Ascend 950 series, which ships from 2026 with Huawei's self-developed HBM, domestic memory a generation or two behind the frontier.
[TrendForce 2026] TrendForce. TSMC sees AI wafer demand rising 11x from 2022–2026, targets CoWoS with 24 HBM stacks in 2029. https://www.trendforce.com/news/2026/05/14/news-tsmc-sees-ai-wafer-demand-rising-11x-from-2022-2026-targets-cowos-with-24-hbm-stacks-in-2029/
[U.S. Bureau of Industry and Security 2026] U.S. Bureau of Industry and Security. Revision to license review policy for advanced computing commodities. https://www.federalregister.gov/documents/2026/01/15/2026-00789/revision-to-license-review-policy-for-advanced-computing-commodities
The cached file is a Federal Register access-blocked page, returning only a CAPTCHA challenge with no retrievable document content.
[Uesato et al. 2022] Uesato, Kushman, Kumar, Song, Siegel, Wang, Creswell, Irving, Higgins. Solving math word problems with process- and outcome-based feedback. arXiv preprint arXiv:2211.14275. https://arxiv.org/abs/2211.14275
This paper compares process-supervised and outcome-supervised reward models for math reasoning on GSM8K, finding that a PRM reduces reasoning trace error from 14.0
[United States District Court for the District of Delaware 2025] United States District Court for the District of Delaware. Thomson reuters enterprise centre GmbH v. ross intelligence inc., no. 1:20-cv-613, memorandum opinion. https://www.ded.uscourts.gov/sites/ded/files/opinions/20-613_5.pdf
[Vals AI 2026] Vals AI. GPQA diamond benchmark leaderboard. https://www.vals.ai/benchmarks/gpqa
[Van Bulck et al. 2018] Van Bulck, Minkin, Weisse, Genkin, Kasikci, Piessens, Silberstein, Wenisch, Yarom, Strackx. Foreshadow: Extracting the keys to the intel SGX kingdom with transient out-of-order execution. https://www.usenix.org/conference/usenixsecurity18/presentation/bulck
The transient-execution attack that extracted SGX's own attestation keys, resetting the field's expectations: a trusted execution boundary is an engineering artifact, not a proof.
[Variety 2024] Variety. News corp inks OpenAI licensing deal potentially worth more than $250 million. https://variety.com/2024/digital/news/news-corp-openai-licensing-deal-1236013734/
[Vaswani et al. 2017] Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin. Attention is all you need. https://arxiv.org/abs/1706.03762
Vaswani et al. introduce the Transformer, a sequence transduction architecture built entirely on multi-head attention with no recurrence or convolutions, achieving state-of-the-art machine translation quality with greater parallelism.
[Vidgen et al. 2024] Vidgen, Agrawal, Ahmed, others. Introducing v0.5 of the AI safety benchmark from mlcommons. arXiv preprint arXiv:2404.12241. https://arxiv.org/abs/2404.12241
MLCommons introduces AI Safety Benchmark v0.5, a preliminary framework with 43,090 test prompts across seven hazard categories to evaluate safety risks of chat-tuned language models.
[Villalobos et al. 2024] Villalobos, Ho, Sevilla, Besiroglu, Heim, Hobbhahn. Will we run out of data? Limits of LLM scaling based on human-generated data. arXiv preprint arXiv:2211.04325. https://arxiv.org/abs/2211.04325
This paper forecasts that LLM training demand will exhaust all available public human-generated text data between 2026 and 2032, and surveys synthetic data and transfer learning as potential mitigations.
[Vincent 2011] Vincent. A connection between score matching and denoising autoencoders. Neural Computation 23(7):1661–1674. https://direct.mit.edu/neco/article-abstract/23/7/1661/7677
[Vipra and Korinek 2023] Vipra, Korinek. Market concentration implications of foundation models. https://arxiv.org/abs/2311.01550
Vipra and Korinek argue that the most capable foundation models tend toward natural monopoly, while behind-frontier models can face intense competition, shaping antitrust and regulatory priorities.
[Visa 2025] Visa. Visa introduces trusted agent protocol: An ecosystem-led framework for AI commerce. https://investor.visa.com/news/news-details/2025/Visa-Introduces-Trusted-Agent-Protocol-An-Ecosystem-Led-Framework-for-AI-Commerce/default.aspx
Visa's agent-identity framework, built with Cloudflare: legitimate agents sign their requests so merchant defenses can recognize known robots instead of fingerprinting for humanity.
[Visa 2025] Visa. Visa and partners complete secure AI transactions, setting the stage for mainstream adoption in 2026. https://usa.visa.com/about-visa/newsroom/press-releases.releaseId.21961.html
[vLLM 2026] vLLM. Batch invariance. https://docs.vllm.ai/en/latest/features/batch_invariance/
vLLM's batch-invariance mode swaps in deterministic kernels so the same request produces the same output regardless of how it is batched, trading throughput for reproducibility.
[vLLM Multimodal Workstream 2025] vLLM Multimodal Workstream. Encoder disaggregation for scalable multimodal model serving. https://vllm.ai/blog/2025-12-15-vllm-epd
vLLM's native encode-prefill-decode (EPD) disaggregation, available since v0.11.1, runs the vision encoder in its own pool and reports goodput roughly doubling on multi-image workloads with 20-50
[W3C 2025] W3C. Verifiable credentials data model v2.0. https://www.w3.org/TR/vc-data-model-2.0/
The W3C Recommendation standardizing cryptographically verifiable claims, the envelope format AP2's mandates travel in, finished four months before the payment protocols shipped.
[Wallace et al. 2024] Wallace, Xiao, Leike, Weng, Heidecke, Beutel. The instruction hierarchy: Training llms to prioritize privileged instructions. arXiv preprint arXiv:2404.13208. https://arxiv.org/abs/2404.13208
Instruction hierarchy trains models to selectively ignore lower-privileged conflicting instructions, improving robustness to prompt injections and jailbreaks with minimal capability degradation.
[Wallace et al. 2024] Wallace, Xiao, Leike, Weng, Heidecke, Beutel. The instruction hierarchy: Training llms to prioritize privileged instructions. arXiv preprint arXiv:2404.13208. https://arxiv.org/abs/2404.13208
This paper proposes an instruction hierarchy that trains LLMs to prioritize system prompts over user and tool messages, substantially improving robustness against prompt injection, jailbreaks, and system message extraction attacks.
[Wan et al. 2025] Wan, Klyman, Kapoor, Maslej, Longpre, Xiong, Liang, Bommasani. The 2025 foundation model transparency index. https://arxiv.org/abs/2512.10169
The 2025 FMTI reports that average developer transparency fell from 58 to 40 out of 100, with training data, training compute, and post-deployment usage among the most opaque areas.
[Wang and Isola 2020] Wang, Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. https://arxiv.org/abs/2005.10242
This paper identifies alignment and uniformity as the two key properties that contrastive representation learning optimizes on the unit hypersphere, and shows directly optimizing quantifiable metrics for both achieves comparable or better downstream task performance.
[Wang and Liu 2021] Wang, Liu. Understanding the behaviour of contrastive loss. https://arxiv.org/abs/2012.09740
This paper analyzes unsupervised contrastive loss as a hardness-aware function and shows that temperature T governs a uniformity-tolerance dilemma, where too-small T enforces uniformity but disrupts semantic structure.
[Wang et al. 2022] Wang, Wei, Schuurmans, Le, Chi, Narang, Chowdhery, Zhou. Self-consistency improves chain of thought reasoning in language models. https://arxiv.org/abs/2203.11171
Self-consistency replaces greedy decoding in chain-of-thought prompting by sampling diverse reasoning paths and selecting the most consistent answer via majority vote, substantially improving reasoning accuracy.
[Wang et al. 2022] Wang, Yang, Huang, Jiao, Yang, Jiang, Majumder, Wei. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533. https://arxiv.org/abs/2212.03533
E5 trains general-purpose text embeddings via contrastive pre-training on CCPairs, a curated 270M web-scale text pair dataset, achieving state-of-the-art results on BEIR and MTEB benchmarks.
[Wang et al. 2023] Wang, Chen, Wu, Zhang, Zhou, Liu, others. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111. https://arxiv.org/abs/2301.02111
VALL-E treats TTS as conditional codec language modeling, training on 60K hours of speech to enable zero-shot personalized speech synthesis from a 3-second acoustic prompt.
[Wang et al. 2023] Wang, Dong, Zeng, Adams, Sreedhar, Egert, Delalleau, Scowcroft, Kant, Swope, Kuchaiev. HelpSteer: Multi-attribute helpfulness dataset for SteerLM. arXiv preprint arXiv:2311.09528. https://arxiv.org/abs/2311.09528
HelpSteer annotates responses on helpfulness plus correctness, coherence, complexity, and verbosity, making preference data more diagnostic than a single scalar label.
[Wang et al. 2023] Wang, Yang, Huang, Yang, Majumder, Wei. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368. https://arxiv.org/abs/2401.00368
E5-Mistral-7B fine-tunes Mistral-7B on LLM-generated synthetic data spanning 93 languages with contrastive loss, achieving state-of-the-art text embedding results on BEIR and MTEB in under 1k training steps.
[Wang et al. 2023] Wang, Li, Chen, Cai, Zhu, Lin, Cao, Liu, Liu, Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926. https://arxiv.org/abs/2305.17926
This paper exposes a positional bias in LLM-as-judge evaluation where swapping response order causes GPT-4 and ChatGPT to reverse their verdicts, and proposes three calibration strategies to mitigate it.
[Wang et al. 2023] Wang, Ma, Dong, Huang, Wang, Ma, Yang, Wang, Wu, Wei. BitNet: Scaling 1-bit transformers for large language models. arXiv preprint arXiv:2310.11453. https://arxiv.org/abs/2310.11453
BitNet proposes a 1-bit Transformer architecture for large language models that trains 1-bit weights from scratch, achieving competitive perplexity while substantially reducing memory and energy consumption, and exhibiting a scaling law similar to FP16 Transformers.
[Wang et al. 2024] Wang, Gao, Zhao, Sun, Dai. Auxiliary-loss-free load balancing strategy for mixture-of-experts. https://arxiv.org/abs/2408.15664
Loss-Free Balancing maintains balanced expert load in MoE models by dynamically updating per-expert routing biases, eliminating auxiliary-loss interference gradients and improving model performance.
[Wang et al. 2024] Wang, Bai, Tan, Wang, Fan, Bai, Chen, Liu, Wang, Ge, Fan, Dang, Du, Ren, Men, Liu, Zhou, Zhou, Lin. Qwen2-VL: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191. https://arxiv.org/abs/2409.12191
Qwen2-VL introduces Naive Dynamic Resolution and M-RoPE to build a vision-language model series (2B, 8B, 72B) that processes images and videos at any resolution, matching GPT-4o on multimodal benchmarks.
[Wang and others 2024] Wang, others. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869. https://arxiv.org/abs/2409.18869
Emu3 trains a single 8B-parameter Transformer on discrete image, video, and text tokens using only next-token prediction, matching or surpassing SDXL and LLaVA-1.6 without diffusion or CLIP.
[Wang et al. 2024] Wang, Chen, Yuan, Zhang, Li, Peng, Ji. Executable code actions elicit better LLM agents. https://arxiv.org/abs/2402.01030
CodeAct has an LLM agent emit executable Python as its action space instead of one structured tool call per turn, composing many tool invocations in a single action and raising success rates by up to 20
[Wang et al. 2024] Wang, Ma, Wei. BitNet a4.8: 4-bit activations for 1-bit llms. arXiv preprint arXiv:2411.04965. https://arxiv.org/abs/2411.04965
BitNet a4.8 introduces a hybrid quantization and sparsification strategy enabling 4-bit activations for 1-bit LLMs, matching BitNet b1.58 performance while activating only 55
[Wang et al. 2024] Wang, Zhou, Song, Mao, Ma, Wang, Xia, Wei. 1-bit AI infra: Part 1.1, fast and lossless BitNet b1.58 inference on cpus. arXiv preprint arXiv:2410.16144. https://arxiv.org/abs/2410.16144
bitnet.cpp is a CPU inference framework for 1-bit LLMs (BitNet b1.58) that achieves 2.37x to 6.17x speedups and up to 82.2
[Wang and others 2025] Wang, others. RAGEN: Understanding self-evolution in LLM agents via multi-turn reinforcement learning. arXiv preprint arXiv:2504.20073. https://arxiv.org/abs/2504.20073
RAGEN introduces StarPO, a multi-turn trajectory-level RL framework for LLM agents, and identifies training instability patterns (Echo Trap) plus rollout design principles that enable stable agent self-evolution.
[Wang et al. 2025] Wang, Gao, Wang, Liu, Sun, Cheng, Shi, Du, Li. MCPTox: a benchmark for tool poisoning attack on real-world MCP servers. https://arxiv.org/abs/2508.14925
MCPTox evaluates tool-poisoning attacks on 45 live MCP servers and 353 tools, showing that malicious instructions hidden in tool metadata can drive unsafe tool use.
[Waxell 2025] Waxell. The $47,000 agent loop: Why token budget alerts aren't budget enforcement. https://dev.to/waxell/the-47000-agent-loop-why-token-budget-alerts-arent-budget-enforcement-389i
[Wei et al. 2022] Wei, Tay, Bommasani, Raffel, Zoph, others. Emergent abilities of large language models. Transactions on Machine Learning Research. https://arxiv.org/abs/2206.07682
This paper defines and surveys emergent abilities of large language models, capabilities absent in smaller models that appear unpredictably past a critical scale threshold, spanning few-shot prompting and augmented prompting tasks.
[Wei et al. 2022] Wei, Wang, Schuurmans, Bosma, Ichter, Xia, Chi, Le, Zhou. Chain-of-thought prompting elicits reasoning in large language models. https://arxiv.org/abs/2201.11903
Chain-of-thought prompting, which adds intermediate reasoning steps as few-shot exemplars, significantly improves large language model performance on arithmetic, commonsense, and symbolic reasoning tasks.
[Wei et al. 2024] Wei, Yang, Song, Lu, Hu, Huang, Tran, Peng, Liu, Huang, Du, Le. Long-form factuality in large language models. arXiv preprint arXiv:2403.18802. https://arxiv.org/abs/2403.18802
SAFE decomposes long-form answers into atomic claims and checks each with an LLM agent that issues Google Search queries, agreeing with crowd annotators 72
[Wei et al. 2024] Wei, Nguyen, Chung, Jiao, Papay, Glaese, Schulman, Fedus. Measuring short-form factuality in large language models. arXiv preprint arXiv:2411.04368. https://arxiv.org/abs/2411.04368
SimpleQA evaluates short fact-seeking questions with single indisputable answers and grades responses as correct, incorrect, or not attempted to measure whether models know what they know.
[Wei et al. 2025] Wei, Duchenne, Copet, Carbonneaux, Zhang, Fried, Synnaeve, Singh, Wang. SWE-RL: Advancing LLM reasoning via reinforcement learning on open software evolution. arXiv preprint arXiv:2502.18449. https://arxiv.org/abs/2502.18449
SWE-RL applies reinforcement learning with rule-based patch-similarity rewards to software evolution data, training Llama3-SWE-RL-70B to achieve 41.0
[Wei et al. 2025] Wei, Sun, Papay, McKinney, Han, Fulford, Chung, Passos, Fedus, Glaese. BrowseComp: a simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516. https://arxiv.org/abs/2504.12516
BrowseComp is a benchmark of 1,266 hard, verifiable web-search questions designed to measure the persistence and creativity of browsing agents, where only OpenAI Deep Research solves roughly half.
[Weiss et al. 2024] Weiss, Ayzenshteyn, Amit, Mirsky. What was your prompt? A remote keylogging attack on AI assistants. https://arxiv.org/abs/2403.09751
Shows that token-length patterns in encrypted streaming responses let a network observer reconstruct 29
[Weller et al. 2024] Weller, Chang, MacAvaney, Lo, Cohan, Van Durme, Lawrie, Soldaini. FollowIR: Evaluating and teaching information retrieval models to follow instructions. arXiv preprint arXiv:2403.15246. https://arxiv.org/abs/2403.15246
FOLLOWIR introduces a benchmark and training set, grounded in TREC annotator narratives, for evaluating and teaching information retrieval models to follow detailed natural-language instructions.
[Wen et al. 2025] Wen, Liu, Zheng, Xu, Ye, Wu, Liang, Wang, Li, Miao, Bian, Yang. Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base llms. arXiv preprint arXiv:2506.14245. https://arxiv.org/abs/2506.14245
This paper argues that ordinary Pass@K can credit correct answers with flawed reasoning and proposes CoT-Pass@K, requiring both reasoning path and final answer to be correct.
[Weng 2023] Weng. LLM powered autonomous agents. https://lilianweng.github.io/posts/2023-06-23-agent/
Lilian Weng's blog post surveys LLM-powered autonomous agent systems, decomposing them into three core components: planning, memory, and tool use.
[Wengert 1964] Wengert. A simple automatic derivative evaluation program. Communications of the ACM. https://dl.acm.org/doi/10.1145/355586.364791
The two-page 1964 paper that introduced the evaluation trace, the linearized list of elementary operations that every tape-based autodiff engine still records.
[Wenzek et al. 2020] Wenzek, Lachaux, Conneau, Chaudhary, Guzmán, Joulin, Grave. CCNet: Extracting high quality monolingual datasets from web crawl data. https://aclanthology.org/2020.lrec-1.494/
CCNet is an automatic pipeline that extracts large, high-quality monolingual datasets from Common Crawl by deduplicating documents, identifying language, and filtering with a Wikipedia-based perplexity model.
[White et al. 2024] White, Dooley, Roberts, Pal, Feuer, Jain, Shwartz-Ziv, Jain, Saifullah, Dey, others. LiveBench: a challenging, contamination-limited LLM benchmark. arXiv preprint arXiv:2406.19314. https://arxiv.org/abs/2406.19314
LiveBench updates questions monthly from recent sources and uses objective automatic grading to reduce both contamination and subjective judge bias.
[White et al. 2024] White, Haddad, Osborne, Liu, Abdelmonsef, Varghese, Le Hors. The model openness framework: Promoting completeness and openness for reproducibility, transparency, and usability in artificial intelligence. https://arxiv.org/abs/2403.13784
The Model Openness Framework classifies AI releases by whether code, data, documentation, and model components are complete and open enough for reproducibility, transparency, and reuse.
[White et al. 2025] White, Skarlinski, Laurent, Bou. About 30% of humanity's last exam chemistry/biology answers are likely wrong. https://www.futurehouse.org/research-announcements/hle-exam
[Wijk and others 2025] Wijk, others. RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. arXiv preprint arXiv:2411.15114. https://arxiv.org/abs/2411.15114
RE-Bench is a seven-task ML research engineering benchmark comparing frontier AI agents to 61 human experts, finding agents score 4x higher at 2-hour budgets but humans outperform them at 8-hour and longer budgets.
[Willard and Louf 2023] Willard, Louf. Efficient guided generation for large language models. arXiv preprint arXiv:2307.09702. https://arxiv.org/abs/2307.09702
This paper reformulates LLM guided generation as finite-state machine (FSM) transitions, enabling O(1) average-cost token masking for regular expressions and context-free grammars via a prebuilt vocabulary index.
[Williams et al. 2009] Williams, Waterman, Patterson. Roofline: An insightful visual performance model for multicore architectures. Communications of the ACM. https://doi.org/10.1145/1498765.1498785
The roofline model: attainable throughput is the minimum of peak arithmetic rate and memory bandwidth times arithmetic intensity, making memory-bound versus compute-bound a picture instead of a debate.
[Willison 2022] Willison. Prompt injection attacks against GPT-3. https://simonwillison.net/2022/Sep/12/prompt-injection/
Simon Willison's 2022 blog post names and defines prompt injection as a security exploit where malicious user input overrides a GPT-3 system prompt, analogous to SQL injection.
[Willison 2025] Willison. Comparing the memory implementations of claude and ChatGPT. https://simonwillison.net/2025/Sep/12/claude-memory/
A dissection of the two shipped designs: one injects a synthesized profile into every conversation automatically, the other starts blank and searches raw history through explicit tools on demand.
[Willison 2025] Willison. The lethal trifecta for AI agents: private data, untrusted content, and external communication. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
Simon Willison's blog post defines the "lethal trifecta" for LLM agents: combining access to private data, exposure to untrusted content, and external communication ability enables prompt injection data-exfiltration attacks.
[Wood Mackenzie 2025] Wood Mackenzie. Power transformers and distribution transformers will face supply deficits of 30% and 10% in 2025. https://www.woodmac.com/press-releases/power-transformers-and-distribution-transformers-will-face-supply-deficits-of-30-and-10-in-2025/
[WorkOS 2025] WorkOS. Maxim fateev on durable execution for AI agents. https://workos.com/blog/maxim-fateev-temporal-durable-execution-ai-agents
A WorkOS interview with Temporal co-founder Maxim Fateev on why durable execution is necessary for reliable AI agent workflows that survive failures.
[Wortsman et al. 2022] Wortsman, Ilharco, Gadre, Roelofs, Gontijo-Lopes, Morcos, Namkoong, Farhadi, Carmon, Kornblith, Schmidt. Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. https://arxiv.org/abs/2203.05482
Averaging the weights of multiple fine-tuned models ("model soups") improves accuracy and out-of-distribution robustness over the best single model without any extra inference cost.
[Xia et al. 2024] Xia, Gao, Zeng, Chen. Sheared LLaMA: Accelerating language model pre-training via structured pruning. https://arxiv.org/abs/2310.06694
Sheared LLaMA uses targeted structured pruning and dynamic batch loading to compress LLaMA2-7B into 1.3B and 2.7B models at 3
[Xiao et al. 2022] Xiao, Lin, Seznec, Wu, Demouth, Han. SmoothQuant: Accurate and efficient post-training quantization for large language models. https://arxiv.org/abs/2211.10438
SmoothQuant enables training-free W8A8 post-training quantization for LLMs by migrating activation outliers to weights via a mathematically equivalent per-channel scaling transformation.
[Xiao et al. 2024] Xiao, Tian, Chen, Han, Lewis. Efficient streaming language models with attention sinks. https://arxiv.org/abs/2309.17453
StreamingLLM enables LLMs to handle infinite-length token sequences without fine-tuning by retaining a few initial attention sink tokens alongside a rolling KV cache window.
[Xiao et al. 2024] Xiao, Liu, Zhang, Muennighoff, Lian, Nie. C-pack: Packed resources for general chinese embeddings. https://arxiv.org/abs/2309.07597
C-Pack releases BGE embedding models, the C-MTP training dataset (100M Chinese text pairs), and the C-MTEB benchmark (35 datasets, 6 tasks) to advance general Chinese text embeddings.
[Xie et al. 2023] Xie, Pham, Dong, Du, Liu, Lu, Liang, Le, Ma, Yu. DoReMi: Optimizing data mixtures speeds up language model pretraining. https://openreview.net/forum?id=lXuByUeHhd
DoReMi uses Group DRO on a small proxy model to find domain mixture weights for pretraining, improving downstream accuracy by 6.5
[Xie et al. 2023] Xie, Kawaguchi, Zhao, Zhao, Kan, He, Xie. Self-evaluation guided beam search for reasoning. arXiv preprint arXiv:2305.00633. https://arxiv.org/abs/2305.00633
This paper proposes a stepwise self-evaluation mechanism integrated with stochastic beam search to guide LLM multi-step reasoning, outperforming Codex baselines by up to 9.56
[Xie et al. 2024] Xie, Zhang, Chen, Li, Zhao, Cao, Hua, Cheng, Shin, Lei, Liu, Xu, Zhou, Savarese, Xiong, Zhong, Yu. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv preprint arXiv:2404.07972. https://arxiv.org/abs/2404.07972
369 tasks on real operating systems with execution-based checking: humans above 72
[Xie et al. 2025] Xie, Mao, Bai, Zhang, Wang, Lin, Gu, Chen, Yang, Shou. Show-o: One single transformer to unify multimodal understanding and generation. https://arxiv.org/abs/2408.12528
Show-o is a single transformer that unifies multimodal understanding and generation by combining autoregressive modeling for text with discrete diffusion for images.
[Xiong et al. 2020] Xiong, Yang, He, Zheng, Zheng, Xing, Zhang, Lan, Wang, Liu. On layer normalization in the transformer architecture. https://arxiv.org/abs/2002.04745
This paper uses mean field theory to show that placing layer normalization inside residual blocks (Pre-LN) yields well-behaved gradients at initialization, allowing Transformer training without learning rate warm-up.
[Xiong et al. 2021] Xiong, Xiong, Li, Tang, Liu, Bennett, Ahmed, Overwijk. Approximate nearest neighbor negative contrastive learning for dense text retrieval. https://arxiv.org/abs/2007.00808
ANCE proposes selecting hard training negatives for dense text retrieval globally from the entire corpus via an asynchronously updated ANN index, nearly matching a BERT cascade pipeline at 100x lower cost.
[XLang Lab 2025] XLang Lab. Introducing osworld-verified. https://xlang.ai/blog/osworld-verified
[Xu et al. 2021] Xu, Lee, Chen, Hechtman, Huang, Joshi, Krikun, Lepikhin, Ly, Maggioni, Pang, Shazeer, Wang, Wang, Wu, Chen. GSPMD: General and scalable parallelization for ML computation graphs. https://arxiv.org/abs/2105.04663
GSPMD is a compiler-based automatic parallelization system that uses tensor sharding annotations to express data parallelism, tensor parallelism, and pipeline parallelism uniformly, achieving 50-62
[Yadav et al. 2023] Yadav, Tam, Choshen, Raffel, Bansal. TIES-merging: Resolving interference when merging models. Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2306.01708
TIES-MERGING is a training-free method that merges multiple fine-tuned models into one by trimming redundant parameters, resolving sign conflicts, and averaging only aligned values.
[Yan et al. 2024] Yan, Gu, Zhu, Ling. Corrective retrieval augmented generation. arXiv preprint arXiv:2401.15884. https://arxiv.org/abs/2401.15884
CRAG proposes a plug-and-play corrective RAG method using a lightweight retrieval evaluator that triggers different knowledge retrieval actions and falls back to web search when retrieved documents are irrelevant.
[Yan et al. 2025] Yan, Mao, Ji, Zhang, Patil, Stoica, Gonzalez. The berkeley function calling leaderboard (BFCL): From tool use to agentic evaluation of large language models. https://openreview.net/forum?id=2GmDdhBdDk
[Yang et al. 2022] Yang, Hu, Babuschkin, Sidor, Liu, Farhi, Ryder, Pachocki, Chen, Gao. Tensor programs V: Tuning large neural networks via zero-shot hyperparameter transfer. arXiv preprint arXiv:2203.03466. https://arxiv.org/abs/2203.03466
Introduces a width parameterization (muP) under which optimal hyperparameters tuned on a small model transfer zero-shot to a much larger one, removing full-scale hyperparameter sweeps.
[Yang et al. 2023] Yang, Chiang, Zheng, Gonzalez, Stoica. Rethinking benchmark and contamination for language models with rephrased samples. https://arxiv.org/abs/2311.04850
This paper shows that rephrased benchmark test samples (paraphrased or translated) bypass n-gram and embedding decontamination, and proposes an LLM-based decontaminator that detects such contamination in pre-training datasets.
[Yang et al. 2023] Yang, Zhang, Li, Zou, Li, Gao. Set-of-mark prompting unleashes extraordinary visual grounding in GPT-4V. arXiv preprint arXiv:2310.11441. https://arxiv.org/abs/2310.11441
Overlay numbered marks on detected screen regions so the model outputs a mark instead of coordinates, the workaround that carried GUI grounding before natively grounded models.
[Yang et al. 2023] Yang, Swope, Gu, Chalamala, Song, Yu, Godil, Prenger, Anandkumar. LeanDojo: Theorem proving with retrieval-augmented language models. arXiv preprint arXiv:2306.15626. https://arxiv.org/abs/2306.15626
LeanDojo releases tools, data, models, and benchmarks for Lean theorem proving, with retrieval-augmented premise selection as a central bottleneck.
[Yang et al. 2024] Yang, Kautz, Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule. arXiv preprint arXiv:2412.06464. https://arxiv.org/abs/2412.06464
Gated DeltaNet combines a gated forgetting mechanism with delta-rule state updates in a linear-attention recurrence, surpassing Mamba2 and DeltaNet on language modeling and long-context tasks.
[Yang et al. 2024] Yang, Teng, Zheng, Ding, Huang, others. CogVideoX: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072. https://arxiv.org/abs/2408.06072
CogVideoX is a text-to-video diffusion transformer using a 3D causal VAE and expert adaptive LayerNorm to generate 10-second, 768x1360 videos at 16 fps with coherent motion.
[Yang et al. 2025] Yang, Yu, Li, Liu, Huang, Huang, Jiang, Tu, Zhang, Zhou, Lin, Dang, Yang, Yu, Li, Sun, Zhu, Men, He, Xu, Yin, Yu, Qiu, Ren, Yang, Li, Xu, Zhang. Qwen2.5-1M technical report. arXiv preprint arXiv:2501.15383. https://arxiv.org/abs/2501.15383
Reports Qwen2.5-1M, using long-data synthesis, progressive long-context pretraining, and multi-stage SFT to reach 1M-token contexts while preserving short-context performance.
[Yang and others 2025] Yang, others. Qwen3 technical report. https://arxiv.org/abs/2505.09388
Qwen3 introduces a unified dense and MoE model family (0.6B to 235B parameters) that integrates thinking and non-thinking modes with a thinking budget mechanism for inference-time compute control.
[Yao et al. 2022] Yao, Zhao, Yu, Du, Shafran, Narasimhan, Cao. ReAct: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629. https://arxiv.org/abs/2210.03629
ReAct interleaves reasoning traces and task-specific actions, letting LLMs update plans with external observations from tools or environments.
[Yao et al. 2023] Yao, Yu, Zhao, Shafran, Griffiths, Cao, Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. https://arxiv.org/abs/2305.10601
Tree of Thoughts (ToT) is a framework that lets LMs explore multiple reasoning paths via tree search with self-evaluation, raising GPT-4's Game of 24 success rate from 4
[Yao et al. 2023] Yao, Zhao, Yu, Du, Shafran, Narasimhan, Cao. ReAct: Synergizing reasoning and acting in language models. https://arxiv.org/abs/2210.03629
ReAct proposes interleaving verbal reasoning traces and environment actions in LLM prompting, reducing hallucination and outperforming reasoning-only or acting-only baselines on QA, fact verification, and decision-making benchmarks.
[Yao et al. 2024] Yao, Shinn, Razavi, Narasimhan. <span class="nocase">τ-bench</span>: a benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045. https://arxiv.org/abs/2406.12045
τ-bench is a benchmark for evaluating language agents on tool use and dynamic user interaction under domain-specific policies, introducing a pass^k metric for reliability, revealing that gpt-4o succeeds on fewer than 50
[Ye et al. 2025] Ye, Xie, others. Dream 7B: Diffusion large language models. arXiv preprint arXiv:2508.15487. https://arxiv.org/abs/2508.15487
Dream 7B is a 7-billion-parameter discrete diffusion language model that uses AR-based initialization and context-adaptive token-level noise rescheduling to match autoregressive LLM performance while enabling arbitrary-order generation and flexible quality-speed trade-offs.
[Ye et al. 2025] Ye, Huang, Xiao, Chern, Xia, Liu. LIMO: Less is more for reasoning. arXiv preprint arXiv:2502.03387. https://arxiv.org/abs/2502.03387
LIMO shows that SFT on only 800 carefully curated examples elicits strong mathematical reasoning in Qwen2.5-32B-Instruct, achieving 63.3
[Yee and Lochmiller 2025] Yee, Lochmiller. How crusoe powers and transforms AI with stranded energy. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-ai-infrastructure-of-the-future
In a McKinsey interview, Crusoe CEO Chase Lochmiller describes the company's energy-first model: it began by placing modular data centers at oil wells to consume otherwise-flared gas, then grew into a vertically integrated builder of large AI campuses sited on stranded and renewable power.
[Yi et al. 2019] Yi, Yang, Hong, Cheng, Heldt, Kumthekar, Zhao, Wei, Chi. Sampling-bias-corrected neural modeling for large corpus item recommendations. https://dl.acm.org/doi/10.1145/3298689.3346996
[Yin et al. 2024] Yin, Gharbi, Zhang, Shechtman, Durand, Freeman, Park. One-step diffusion with distribution matching distillation. https://arxiv.org/abs/2311.18828
DMD distills a pretrained diffusion model into a one-step image generator by minimizing an approximate KL divergence between real and fake distributions, achieving 2.62 FID on ImageNet 64x64 at 20 FPS with FP16 inference.
[Yin et al. 2024] Yin, Zheng, others. Fast JSON decoding for local llms with compressed finite state machine. https://www.lmsys.org/blog/2024-02-05-compressed-fsm/
SGLang's jump-forward decoding compresses FSM singular paths to prefill multiple tokens at once during constrained decoding, cutting latency by up to 2x and boosting throughput by up to 2.5x.
[Yin et al. 2024] Yin, Zhang, Zhang, Freeman, Durand, Shechtman, Huang. From slow bidirectional to fast autoregressive video diffusion models. arXiv preprint arXiv:2412.07772. https://arxiv.org/abs/2412.07772
Distills a fifty-step bidirectional video diffusion model into a four-step causal autoregressive generator that streams video with a KV cache, importing LLM serving economics into video.
[Yu et al. 2022] Yu, Li, Koh, Zhang, Pang, Qin, Ku, Xu, Baldridge, Wu. Vector-quantized image modeling with improved VQGAN. https://arxiv.org/abs/2110.04627
ViT-VQGAN replaces CNN encoders with Vision Transformers in VQGAN and introduces factorized, L2-normalized codebook learning, improving image generation FID and unsupervised linear-probe accuracy on ImageNet.
[Yu et al. 2022] Yu, Xu, Koh, Luong, others. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789. https://arxiv.org/abs/2206.10789
Parti scales an autoregressive encoder-decoder Transformer for text-to-image generation to 20B parameters, treating image tokens from ViT-VQGAN as a sequence-to-sequence target and achieving state-of-the-art FID on MS-COCO.
[Yu et al. 2022] Yu, Jeong, Kim, Kim, Chun. Orca: a distributed serving system for transformer-based generative models. https://www.usenix.org/conference/osdi22/presentation/yu
[Yu et al. 2024] Yu, Yu, Yu, Huang, Li. Language models are super mario: Absorbing abilities from homologous models as a free lunch. https://arxiv.org/abs/2311.03099
DARE randomly drops and rescales SFT delta parameters to enable merging multiple task-specific language models into one without retraining or GPUs.
[Yu et al. 2025] Yu, Zhang, others. DAPO: An open-source LLM reinforcement learning system at scale. arXiv preprint arXiv:2503.14476. https://arxiv.org/abs/2503.14476
DAPO proposes Decoupled Clip and Dynamic Sampling Policy Optimization with four techniques (Clip-Higher, Dynamic Sampling, Token-Level Policy Gradient Loss, Overlong Reward Shaping) to enable reproducible large-scale RL training, achieving 50 points on AIME 2024 with Qwen2.5-32B.
[Yuan et al. 2023] Yuan, Yuan, Li, Dong, Lu, Tan, Zhou, Zhou. Scaling relationship on learning mathematical reasoning with large language models. arXiv preprint arXiv:2308.01825. https://arxiv.org/abs/2308.01825
This paper studies scaling relationships for mathematical reasoning and introduces rejection sampling fine-tuning, which augments training data with verifier-selected correct reasoning paths.
[Yuan et al. 2024] Yuan, Shang, Zhou, Dong, Zhou, Xue, Wu, Li, Gu, Lee, Yan, Chen, Sun, Keutzer. LLM inference unveiled: Survey and roofline model insights. arXiv preprint arXiv:2402.16363. https://arxiv.org/abs/2402.16363
This survey analyzes LLM inference efficiency using a Roofline model framework and introduces LLM-Viewer, a tool that identifies memory and compute bottlenecks when deploying LLMs on hardware.
[Yuan et al. 2024] Yuan, Shang, Zhou, Dong, Zhou, Xue, Wu, Li, Gu, Lee, Yan, Chen, Sun, Keutzer. LLM inference unveiled: Survey and roofline model insights. arXiv preprint arXiv:2402.16363. https://arxiv.org/abs/2402.16363
This survey introduces LLM-Viewer, a Roofline model-based framework for systematically analyzing memory and compute bottlenecks in LLM inference, alongside a review of compression, decoding, and system-level optimizations.
[Yuan et al. 2025] Yuan, Gao, Dai, Luo, Zhao, others. Native sparse attention: Hardware-aligned and natively trainable sparse attention. arXiv preprint arXiv:2502.11089. https://arxiv.org/abs/2502.11089
NSA makes sparse attention end-to-end trainable rather than a post-hoc mask: a hierarchical compress-select-window design that speeds up long-context training and decoding while matching full-attention quality.
[Yuan et al. 2025] Yuan, Sriskandarajah, Brakman, Helyar, Beutel, Vallone, Jain. From hard refusals to safe-completions: Toward output-centric safety training. arXiv preprint arXiv:2508.09224. https://arxiv.org/abs/2508.09224
Safe-completions training replaces binary refusal with an output-centric objective that maximizes helpfulness subject to output-safety constraints; shipped in GPT-5, it improves both safety and helpfulness on dual-use prompts where intent classification fails.
[Yuan et al. 2025] Yuan, Xiao, Tao, Wang, Gao, Ding, Xu. Incentivizing reasoning from weak supervision. arXiv preprint arXiv:2505.20072. https://arxiv.org/abs/2505.20072
This paper studies weak-to-strong reasoning supervision and reports that weaker reasoners can recover most of the gains of expensive RL when supervising stronger students.
[Yue et al. 2023] Yue, Ni, Zhang, Zheng, Liu, Zhang, Stevens, Jiang, Ren, Sun, others. MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. arXiv preprint arXiv:2311.16502. https://arxiv.org/abs/2311.16502
MMMU evaluates multimodal models on 11.5K college-level questions spanning six disciplines, requiring domain knowledge, visual understanding, and deliberate reasoning.
[Yue et al. 2024] Yue, Ni, Zhang, Zheng, Liu, Zhang, Stevens, Jiang, Ren, Sun, others. MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. https://arxiv.org/abs/2311.16502
MMMU is a benchmark of 11.5K college-level multimodal questions across 30 subjects and 6 disciplines to evaluate expert-level perception, knowledge, and reasoning in large multimodal models.
[Yue et al. 2024] Yue, Zheng, Ni, Wang, Zhang, Tong, Sun, Yu, Zhang, Sun, others. MMMU-pro: a more robust multi-discipline multimodal understanding benchmark. arXiv preprint arXiv:2409.02813. https://arxiv.org/abs/2409.02813
[Yue et al. 2025] Yue, Chen, Lu, Zhao, Wang, Yue, Song, Huang. Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?. arXiv preprint arXiv:2504.13837. https://arxiv.org/abs/2504.13837
RLVR improves sampling efficiency toward correct answers at low pass@k but does not expand reasoning beyond the base model's existing distribution, which bounds RLVR-trained LLMs.
[Zadouri et al. 2026] Zadouri, Hoehnerbach, Shah, Liu, Thakkar, Dao. FlashAttention-4: Algorithm and kernel pipelining co-design for asymmetric hardware scaling. https://arxiv.org/abs/2603.05451
FlashAttention-4 targets Blackwell with a CuTe-DSL (Python-embedded) implementation, software-emulated exponentials on FMA units, and conditional softmax rescaling, reaching 1.1-1.3x over cuDNN attention on B200.
[Zeff 2025] Zeff. Silicon Valley bets big on 'environments' to train AI agents. https://techcrunch.com/2025/09/21/silicon-valley-bets-big-on-environments-to-train-ai-agents/
A TechCrunch report on startups building RL training environments for AI agents, described as Silicon Valley's next investment wave.
[Zeghidour et al. 2021] Zeghidour, Luebs, Omran, Skoglund, Tagliasacchi. SoundStream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30. https://arxiv.org/abs/2107.03312
SoundStream is an end-to-end neural audio codec with a convolutional encoder/decoder and residual vector quantizer (RVQ) that outperforms Opus and EVS at low bitrates (3-18 kbps) on speech, music, and general audio.
[Zelikman et al. 2022] Zelikman, Wu, Mu, Goodman. STaR: Bootstrapping reasoning with reasoning. arXiv preprint arXiv:2203.14465. https://arxiv.org/abs/2203.14465
STaR iteratively generates rationales, keeps those that lead to correct answers, and fine-tunes on the resulting traces to bootstrap reasoning from a small rationale seed set.
[Zeng et al. 2024] Zeng, Liu, Mullins, Peran, Fernandez, Harkous, Narasimhan, Proud, Kumar, Radharapu, Sturman, Wahltinez. ShieldGemma: Generative AI content moderation based on gemma. arXiv preprint arXiv:2407.21772. https://arxiv.org/abs/2407.21772
ShieldGemma is a suite of LLM-based content moderation models (2B to 27B) built on Gemma2 that classify harmful content across six harm types in both user inputs and model outputs.
[Zep 2025] Zep. Zep: Temporal knowledge graph for agent memory. https://www.getzep.com/
Zep is an enterprise agent memory platform that stores facts in a temporal context graph (Graphiti), tracks how facts change over time, and serves retrieved context to agents in under 200 ms.
[Zhai et al. 2023] Zhai, Mustafa, Kolesnikov, Beyer. Sigmoid loss for language image pre-training. https://arxiv.org/abs/2303.15343
SigLIP replaces softmax contrastive loss with a pairwise sigmoid loss for language-image pre-training, enabling memory-efficient large-batch training and better performance at small batch sizes.
[Zhai et al. 2024] Zhai, Liao, Liu, Wang, Li, Cao, Gao, Gong, Gu, He, Lu, Shi. Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations. https://arxiv.org/abs/2402.17152
Meta's HSTU reformulates ranking and retrieval as sequential transduction, deploys trillion-parameter generative recommenders with double-digit online A/B gains, and reports quality scaling as a power law of training compute.
[Zhai and others 2026] Zhai, others. HLE-verified: a systematic verification and structured revision of humanity's last exam. arXiv preprint arXiv:2602.13964. https://arxiv.org/abs/2602.13964
[Zhan et al. 2024] Zhan, Liang, Ying, Kang. InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents. https://aclanthology.org/2024.findings-acl.624/
INJECAGENT is a benchmark of 1,054 test cases evaluating LLM agent vulnerability to indirect prompt injection attacks, finding ReAct-prompted GPT-4 susceptible 24
[Zhang and Sennrich 2019] Zhang, Sennrich. Root mean square layer normalization. https://arxiv.org/abs/1910.07467
RMSNorm replaces LayerNorm's mean-and-variance normalization with RMS-only scaling, achieving comparable accuracy while reducing per-step runtime by 7 to 64 percent.
[Zhang et al. 2023] Zhang, Sheng, Zhou, Chen, Zheng, Cai, Song, Tian, Ré, Barrett, Wang, Chen. H2O: Heavy-hitter oracle for efficient generative inference of large language models. https://arxiv.org/abs/2306.14048
H2O proposes a KV cache eviction policy that retains "heavy hitter" tokens identified by accumulated attention scores, reducing KV cache memory while preserving LLM generation quality.
[Zhang et al. 2024] Zhang, Zhou, Yang, Wang, Liu, others. SpeechTokenizer: Unified speech tokenizer for speech language models. https://arxiv.org/abs/2308.16692
SpeechTokenizer is an RVQ-based unified speech tokenizer for speech LMs that disentangles semantic content and acoustic details across hierarchical quantizer layers via semantic distillation from HuBERT.
[Zhang et al. 2024] Zhang, Hosseini, Bansal, Kazemi, Kumar, Agarwal. Generative verifiers: Reward modeling as next-token prediction. arXiv preprint arXiv:2408.15240. https://arxiv.org/abs/2408.15240
GenRM trains LLM verifiers with next-token prediction rather than discriminative classification, enabling verifier chain-of-thought and test-time voting for Best-of-N selection.
[Zhang et al. 2025] Zhang, Zheng, Wu, Zhang, Lin, others. The lessons of developing process reward models in mathematical reasoning. arXiv preprint arXiv:2501.07301. https://arxiv.org/abs/2501.07301
Monte-Carlo-estimated step labels yield weaker PRMs than LLM-as-judge and human annotation; a consensus filter that keeps only steps where both agree produces the stronger Qwen2.5-Math-PRM, evaluated on the released ProcessBench.
[Zhang et al. 2025] Zhang, Geng, Yu, Yin, Zhang, others. The landscape of agentic reinforcement learning for llms: a survey. arXiv preprint arXiv:2509.02547. https://arxiv.org/abs/2509.02547
A survey of agentic reinforcement learning for LLMs, formalizing the shift from single-step RLHF to multi-step POMDP-based agents with planning, tool use, memory, and self-improvement capabilities.
[Zhang et al. 2025] Zhang, Li, Long, Zhang, Lin, Yang, Xie, Yang, Liu, Lin, Huang, Zhou. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176. https://arxiv.org/abs/2506.05176
Qwen3 Embedding introduces text embedding and reranking models (0.6B, 4B, 8B) built on Qwen3 foundation models with multi-stage training and LLM-synthesized data, achieving state-of-the-art results on MTEB.
[Zhao et al. 2023] Zhao, Gu, Varma, Luo, Huang, Xu, Wright, Shojanazeri, Ott, Shleifer, Desmaison, Balioglu, Damania, Nguyen, Chauhan, Hao, Mathews, Li. PyTorch FSDP: Experiences on scaling fully sharded data parallel. https://arxiv.org/abs/2304.11277
PyTorch FSDP is an industry-grade fully sharded data parallel training system that shards model parameters across GPUs, achieving near-linear TFLOPS scalability for large models while matching DDP performance on small ones.
[Zhao et al. 2025] Zhao, Wu, Yue, others. Absolute zero: Reinforced self-play reasoning with zero data. arXiv preprint arXiv:2505.03335. https://arxiv.org/abs/2505.03335
Absolute Zero trains a single model to propose tasks that maximize its own learning progress and to solve them, with a code executor verifying both roles, reaching strong coding and math reasoning with no external data.
[Zheng et al. 2021] Zheng, Han, Polu. MiniF2F: a cross-system benchmark for formal olympiad-level mathematics. arXiv preprint arXiv:2109.00110. https://arxiv.org/abs/2109.00110
MiniF2F provides 488 formal Olympiad-level problem statements across Metamath, Lean, Isabelle, and HOL Light to benchmark neural theorem proving.
[Zheng et al. 2023] Zheng, Yin, Xie, Sun, Huang, Yu, Cao, Kozyrakis, Stoica, Gonzalez, Barrett, Sheng. SGLang: Efficient execution of structured language model programs. https://arxiv.org/abs/2312.07104
SGLang is a system combining a structured generation language and a runtime with RadixAttention for KV cache reuse and compressed FSM-based constrained decoding, achieving up to 6.4x higher throughput for complex LLM programs.
[Zheng et al. 2023] Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica. Judging LLM-as-a-judge with MT-bench and chatbot arena. arXiv preprint arXiv:2306.05685. https://arxiv.org/abs/2306.05685
This paper introduces MT-bench and Chatbot Arena to validate LLM-as-a-judge, showing GPT-4 matches human preference ratings at over 80
[Zheng et al. 2024] Zheng, Yin, Xie, Sun, Huang, Yu, Cao, Kozyrakis, Stoica, Gonzalez, Barrett, Sheng. SGLang: Efficient execution of structured language model programs. https://arxiv.org/abs/2312.07104
SGLang introduces RadixAttention for automatic KV cache reuse and compressed finite-state machines for faster constrained decoding, achieving up to 6.4x throughput gains over vLLM and similar systems.
[Zheng et al. 2025] Zheng, Liu, Yu, Men, Yang, Zhou, Lin, others. Group sequence policy optimization. arXiv preprint arXiv:2507.18071. https://arxiv.org/abs/2507.18071
GSPO replaces GRPO's token-level importance ratios with sequence-level ratios and clipping, stabilizing MoE RL training and improving efficiency in Qwen3 models.
[Zheng et al. 2025] Zheng, Zhao, Chen. Prosperity before collapse: How far can off-policy RL reach with stale data on llms?. arXiv preprint arXiv:2510.01161. https://arxiv.org/abs/2510.01161
M2PO constrains the second moment of importance weights to enable stable off-policy RL training on LLMs with data stale by up to 256 model updates, matching on-policy GRPO performance.
[Zhong et al. 2023] Zhong, Guo, Gao, Ye, Wang. MemoryBank: Enhancing large language models with long-term memory. arXiv preprint arXiv:2305.10250. https://arxiv.org/abs/2305.10250
Long-term memory for companion assistants with memory strength updated on an Ebbinghaus-style forgetting curve, the earliest serious treatment of decay as a design element.
[Zhong et al. 2024] Zhong, Liu, Chen, Hu, Zhu, Liu, Jin, Zhang. DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. https://arxiv.org/abs/2401.09670
DistServe improves LLM serving goodput by disaggregating prefill and decoding onto separate GPUs, eliminating interference and enabling independent resource and parallelism optimization for TTFT and TPOT.
[Zhou et al. 2020] Zhou, Gu, Neubig. Understanding knowledge distillation in non-autoregressive machine translation. https://arxiv.org/abs/1911.02727
This paper explains why sequence-level knowledge distillation helps non-autoregressive machine translation (NAT) by showing that distillation reduces training data complexity, and that optimal data complexity correlates with NAT model capacity.
[Zhou et al. 2022] Zhou, Lei, Liu, Du, Huang, Zhao, Dai, Chen, Le, Laudon. Mixture-of-experts with expert choice routing. https://arxiv.org/abs/2202.09368
Expert Choice MoE proposes letting each expert select its top-k tokens instead of each token choosing experts, guaranteeing perfect load balancing and achieving over 2x faster training convergence than Switch Transformer and GShard.
[Zhou et al. 2022] Zhou, Schärli, Hou, Wei, Scales, Wang, Schuurmans, Cui, Bousquet, Le, Chi. Least-to-most prompting enables complex reasoning in large language models. https://arxiv.org/abs/2205.10625
Least-to-most prompting decomposes a complex problem into simpler subproblems solved sequentially, enabling large language models to generalize to harder problems than those in the prompt exemplars.
[Zhou et al. 2023] Zhou, Liu, Xu, Iyer, Sun, Mao, Ma, Efrat, Yu, Yu, Zhang, Ghosh, Lewis, Zettlemoyer, Levy. LIMA: Less is more for alignment. arXiv preprint arXiv:2305.11206. https://arxiv.org/abs/2305.11206
LIMA shows that fine-tuning a 65B LLaMa model on only 1,000 carefully curated prompt-response pairs, without RLHF, produces strong alignment, supporting the Superficial Alignment Hypothesis.
[Zhou et al. 2023] Zhou, Xu, Zhu, Zhou, Lo, Sridhar, Cheng, Ou, Bisk, Fried, Alon, Neubig. WebArena: a realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854. https://arxiv.org/abs/2307.13854
Self-hosted realistic web tasks (e-commerce, forums, code hosting) where the best GPT-4 agent completed 14.41
[Zhou et al. 2025] Zhou, Yu, Babu, Tirumala, Yasunaga, Shamis, Kahn, Ma, Zettlemoyer, Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. https://arxiv.org/abs/2408.11039
Transfusion trains a single transformer on mixed text and image data by applying next-token prediction loss to text and DDPM diffusion loss to images, scaling to 7B parameters with quality matching dedicated models.
[Zhou and others 2026] Zhou, others. Externalization in LLM agents: a unified review of memory, skills, protocols and harness engineering. https://arxiv.org/abs/2604.08224
This survey proposes externalization as the unifying principle of LLM agent design, framing memory, skills, and protocols as cognitive artifacts that offload model burdens into persistent external infrastructure coordinated by a harness layer.
[Zhu et al. 2024] Zhu, Yin, Deng, Almeida, Zhou. Confidential computing on NVIDIA hopper gpus: a performance benchmark study. arXiv preprint arXiv:2409.03992. https://arxiv.org/abs/2409.03992
Benchmarks H100 confidential-computing mode on LLM serving: under 5
[Zhu et al. 2025] Zhu, You, Xing, Huang, others. LLaDA-MoE: a sparse MoE diffusion language model. arXiv preprint arXiv:2509.24389. https://arxiv.org/abs/2509.24389
[Zhu and Li 2025] Zhu, Li. Towards concise and adaptive thinking in large reasoning models: a survey. arXiv preprint arXiv:2507.09662. https://arxiv.org/abs/2507.09662
This survey reviews methods, benchmarks, and open problems for concise and adaptive thinking in reasoning models, addressing the cost of unnecessarily long chains.
[Zoph et al. 2022] Zoph, Bello, Kumar, Du, Huang, Dean, Shazeer, Fedus. ST-MoE: Designing stable and transferable sparse expert models. https://arxiv.org/abs/2202.08906
ST-MoE-32B is a 269B sparse MoE model that resolves MoE training instability and fine-tuning transfer gaps, achieving state-of-the-art results across diverse NLP benchmarks.
[Zou et al. 2023] Zou, Wang, Carlini, Nasr, Kolter, Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. https://arxiv.org/abs/2307.15043
GCG introduces an automated greedy coordinate gradient method to find adversarial suffixes that jailbreak aligned LLMs, transferring to black-box models including ChatGPT, Bard, and Claude.
[Zou et al. 2024] Zou, Phan, Wang, Duenas, Lin, Andriushchenko, Wang, Kolter, Fredrikson, Hendrycks. Improving alignment and robustness with circuit breakers. arXiv preprint arXiv:2406.04313. https://arxiv.org/abs/2406.04313
Circuit breakers use representation rerouting to redirect internal model representations away from harmful outputs, achieving adversarial robustness across unseen attacks without sacrificing capability in LLMs, multimodal models, and agents.

Comments

Log in to comment