Skimmaxxer

30 papers·7 read in full

The library

Every paper the project has opened, and how far each one was taken.

1655 concepts·297 figures and tables·894 connections

How to use this

  1. Pick a paper below.

    Anything under Read in full has been taken apart completely. The rest were opened for one mechanism and closed again.

  2. Choose a way in. Either button on the card.

    • Original source with annotationsThe paper as it was printed. The concepts behind whatever paragraph you are on sit beside it, so a term that arrives undefined is explained without leaving the page.
    • Paper wikiThe paper reconstructed for skimming. Retold start to finish, with every term on a page of its own and every chapter able to reopen one level deeper.

    You can switch between them at any point without losing your place.

  3. Follow anything you do not know.

    Every term the paper leans on is a page of its own, and the terms inside that page are pages too. Every claim links to the page of the PDF it came from, so you can always check it.

  4. Not on the shelf? Ask for it.

    Use the form at the foot of this page. Papers are added by hand at the moment, and the aim is to have one read within 24 hours.

All of it is written by a large language model. Not by hand, and not peer-reviewed. Read it as a way into a paper rather than as a substitute for one.

Read in full7

1706.03762 arXiv

Attention Is All You Need

Vaswani, Shazeer, Parmar et al. (2017)

  1. L09
  2. L138
  3. L238
  4. L333

gpt-1 openai.com

Improving Language Understanding by Generative Pre-Training

Radford, Narasimhan, Salimans et al. (2018)

  1. L09
  2. L142
  3. L238
  4. L318

120 concepts, grouped into 8 themes that follow the paper's own order.

1810.04805 arXiv

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee et al.

  1. L09
  2. L144
  3. L239
  4. L324

monosemanticity transformer-circuits.pub

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Trenton Bricken, Adly Templeton, Joshua Batson et al.

  1. L09
  2. L144
  3. L273
  4. L363

2502.17424 arXiv

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

Jan Betley, Daniel Tan, Niels Warncke et al.

  1. L09
  2. L175
  3. L241
  4. L313

256 concepts, grouped into 8 themes that follow the paper's own order.

global-workspace transformer-circuits.pub

Verbalizable Representations Form a Global Workspace in Language Models

Wes Gurnee, Nicholas Sofroniew, Adam Pearce et al.

  1. L09
  2. L144
  3. L279
  4. L339

hugging-face-incident metr.org

Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

Hjalmar Wijk, Ajeya Cotra (METR), Ryan Greenblatt (Redwood Research et al.

  1. L09
  2. L145
  3. L283
  4. L343

Up for a full read23

Ask for the one you want read in full.

Each of these was opened for one mechanism and closed again — the concepts listed under it are all that was taken. Any of them could be read end to end next, and what gets read next is decided by what people ask for.

Propose a paper instead →
  1. 1412.7449

    Grammar as a Foreign Language

    Vinyals, Kaiser, Koo et al. (2015)

    read for Attention Is All You Need

  2. 1506.06724

    Aligning Books and Movies (BookCorpus)

    Yukun Zhu, Ryan Kiros, Richard Zemel et al.

    read for BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

  3. 1508.07909

    Neural Machine Translation of Rare Words with Subword Units

    Sennrich, Haddow & Birch (2015)

    read for Attention Is All You Need·Improving Language Understanding by Generative Pre-Training

  4. 1509.06664

    Reasoning about Entailment with Neural Attention

    Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann et al.

    read for Improving Language Understanding by Generative Pre-Training

  5. 1512.00567

    Rethinking the Inception Architecture for Computer Vision

    Szegedy, Vanhoucke, Ioffe et al. (2015)

    read for Attention Is All You Need

  6. 1606.08415

    Gaussian Error Linear Units (GELUs)

    Dan Hendrycks, Kevin Gimpel

    read for Improving Language Understanding by Generative Pre-Training·BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

  7. 1607.06450

    Layer Normalization

    Ba, Kiros & Hinton (2016)

    read for Attention Is All You Need·Improving Language Understanding by Generative Pre-Training

  8. 1608.05859

    Using the Output Embedding to Improve Language Models

    Press & Wolf (2016)

    read for Attention Is All You Need

  9. 1609.08144

    Google's Neural Machine Translation System

    Wu, Schuster, Chen et al. (2016)

    read for Attention Is All You Need·BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

  10. 1705.03551

    TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

    Mandar Joshi, Eunsol Choi, Daniel S. Weld et al.

    read for BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

  11. 1711.05101

    Decoupled Weight Decay Regularization

    Ilya Loshchilov, Frank Hutter

    read for Improving Language Understanding by Generative Pre-Training

  12. 1801.10198

    Generating Wikipedia by Summarizing Long Sequences

    Peter J. Liu, Mohammad Saleh, Etienne Pot et al.

    read for Improving Language Understanding by Generative Pre-Training

  13. 2209.10652

    Toy Models of Superposition

    Nelson Elhage, Tristan Hume, Catherine Olsson et al.

    read for Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

  14. 2209.15189

    Learning by Distilling Context

    Charlie Snell, Dan Klein, Ruiqi Zhong

    read for Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

  15. 2303.08112

    Eliciting Latent Predictions from Transformers with the Tuned Lens

    Nora Belrose, Igor Ostrovsky, Lev McKinney et al.

    read for Verbalizable Representations Form a Global Workspace in Language Models

  16. 2304.03279

    Do the Rewards Justify the Means? The MACHIAVELLI Benchmark

    Alexander Pan, Jun Shern Chan, Andy Zou et al.

    read for Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

  17. 2308.09124

    Linearity of Relation Decoding in Transformer Language Models

    Evan Hernandez, Arnab Sen Sharma, Tal Haklay et al.

    read for Verbalizable Representations Form a Global Workspace in Language Models

  18. 2312.03732

    A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA

    Damjan Kalajdzievski (Tenyx)

    read for Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

  19. 2401.05566

    Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training

    Evan Hubinger, Carson Denison, Jesse Mu et al.

    read for Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

  20. 2402.10260

    A StrongREJECT for Empty Jailbreaks

    Alexandra Souly, Qingyuan Lu, Dillon Bowen et al.

    read for Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

  21. 2408.02946

    Scaling Trends for Data Poisoning in LLMs

    Dillon Bowen, Brendan Murphy, Will Cai et al.

    read for Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

  22. 2506.02548

    CyberGym: Evaluating AI Agents' Cybersecurity Capabilities at Scale

    Wang, Shi, He et al. (2026)

    read for Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

  23. 2511.18397

    Natural Emergent Misalignment from Reward Hacking in Production RL

    Monte MacDiarmid, Benjamin Wright, Jonathan Uesato et al.

    read for Verbalizable Representations Form a Global Workspace in Language Models

Propose a paper

Anything that is not on the shelf. Send the link and say what you want out of it — what a read would have to explain for the paper to land.

An address is only ever used to tell you the read is up.