---
title: "arXiv has reached lift-off. Is it acceleration or slop?"
subtitle: "Where the surge in preprints is coming from, and what the metadata can and cannot tell us"
author: "Trey Saddler"
date: 2026-10-01
categories: [arXiv, AI, Data Analysis]
jupyter: python3
format:
html:
toc: true
toc-depth: 2
code-fold: true
code-summary: "Show the code"
code-tools: true
fig-format: svg
fig-align: center
df-print: kable
include-in-header: arxiv-liftoff/assets/view-toggle.html
execute:
warning: false
---
A post on LinkedIn recently put the question plainly: monthly submissions to arXiv have "reached lift-off", and either AI has meaningfully accelerated research, or preprints are filling up with low-effort slop. How would we measure which is true? In the comments, Srijit Seal suggested a place to start: work out where the growth is coming from and how fast it is expanding, and check whether credible venues show the same explosion. If they do not, look for mills that "just submit papers every hour".
This article does that analysis with public data. Use the switch under the title to move between the takeaways and the full walkthrough, which shows how each dataset was fetched, every processing step, and the code behind every number.
::: {.thorough}
The analysis is a single Quarto document running on the Jupyter kernel. This first cell sets up paths and one shared chart style.
:::
```{python}
#| label: setup
from pathlib import Path
import json, re, time, zipfile
import matplotlib.pyplot as plt
import matplotlib.ticker as mtick
import numpy as np
import orjson
import pandas as pd
import pyarrow.parquet as pq
import requests
DATA = Path("arxiv-liftoff/data") # relative to posts/; not committed to the site repo
RAW, PROC = DATA / "raw", DATA / "processed"
PARTS = PROC / "parts"
for d in (RAW, PROC, PARTS):
d.mkdir(parents=True, exist_ok=True)
UA = {"User-Agent": "arxiv-liftoff-analysis (mailto:treyosaddler@gmail.com)"}
# One visual system for every chart: thin marks, recessive axes, a fixed colour order.
BLUE, ORANGE, AQUA, YELLOW = "#2a78d6", "#eb6834", "#1baf7a", "#eda100"
INK, INK2, MUTED, GRID = "#0b0b0b", "#52514e", "#898781", "#e1e0d9"
plt.rcParams.update({
"figure.figsize": (7.6, 3.9), "figure.facecolor": "#fcfcfb", "axes.facecolor": "#fcfcfb",
"axes.spines.top": False, "axes.spines.right": False, "axes.spines.left": False,
"axes.edgecolor": "#c3c2b7", "axes.grid": True, "axes.grid.axis": "y", "grid.color": GRID,
"grid.linewidth": 0.8, "axes.axisbelow": True, "xtick.color": MUTED, "ytick.color": MUTED,
"xtick.labelcolor": INK2, "ytick.labelcolor": INK2, "ytick.left": False, "font.size": 10,
"axes.titlesize": 11, "axes.titleweight": "bold", "axes.titlelocation": "left",
"axes.labelcolor": INK2, "text.color": INK, "lines.linewidth": 2, "legend.frameon": False,
"svg.fonttype": "none",
})
thousands = mtick.FuncFormatter(lambda v, _: f"{v:,.0f}")
pct0 = mtick.PercentFormatter(1, decimals=0)
```
## The data {.thorough}
Four public sources, each cached under `data/raw/` the first time it is fetched so that re-rendering does not hit anyone's servers again.
| Source | What it gives us | How it is fetched |
|---|---|---|
| [arXiv monthly submission statistics](https://arxiv.org/stats/monthly_submissions) | Official count of new submissions per month since 1991 | One CSV download |
| [arXiv metadata snapshot on Kaggle](https://www.kaggle.com/datasets/Cornell-University/arxiv) | One JSON record per paper: submitter, authors, categories, versions, comments, abstract | One 1.8 GB zip (5.6 GB of JSON lines) |
| [bioRxiv](https://api.biorxiv.org/) and medRxiv summary APIs | New preprints per month on the two largest life-science servers | One JSON call each |
| [Crossref REST API](https://api.crossref.org/) | Number of journal articles registered, by publication month | One count query per month |
### Step 1: the small sources
```{python}
#| label: fetch-small
def fetch(url, name):
"""Download `url` to data/raw/`name` once; later runs reuse the cached copy."""
path = RAW / name
if not path.exists():
r = requests.get(url, headers=UA, timeout=120)
r.raise_for_status()
path.write_bytes(r.content)
return path
official = pd.read_csv(fetch("https://arxiv.org/stats/get_monthly_submissions", "arxiv_monthly_submissions.csv"))
official["month"] = pd.PeriodIndex(official["month"], freq="M")
official = official.set_index("month")["submissions"]
def rxiv_monthly(server):
j = json.loads(fetch(f"https://api.{server}.org/sum/m", f"{server}_monthly.json").read_text())
key = next(k for k in j if "statistics" in k)
d = pd.DataFrame(j[key])
return pd.Series(d["new_papers"].astype(int).values, index=pd.PeriodIndex(d["month"], freq="M"), name=server)
biorxiv, medrxiv = rxiv_monthly("biorxiv"), rxiv_monthly("medrxiv")
official.tail(4).to_frame()
```
The official arXiv series ends with a partial month (the file was pulled on 1 October 2026), so everything below stops at the last complete month.
### Step 2: journal output from Crossref
Crossref is the DOI registry most journals use. Asking for zero rows with a filter returns just the count of matching records, so one polite request per month gives a monthly series of journal articles by publication date. The loop is resumable: each answer is written to the CSV straight away.
```{python}
#| label: fetch-crossref
def crossref_monthly(kind="journal-article", start="2015-01", end="2026-09"):
path = RAW / f"crossref_{kind}_monthly.csv"
have = pd.read_csv(path) if path.exists() else pd.DataFrame(columns=["month", "count"])
rows, seen = have.to_dict("records"), set(have["month"])
for m in pd.period_range(start, end, freq="M"):
if str(m) in seen:
continue
flt = f"type:{kind},from-pub-date:{m.start_time:%Y-%m-%d},until-pub-date:{m.end_time:%Y-%m-%d}"
r = requests.get("https://api.crossref.org/works", params={"rows": 0, "filter": flt}, headers=UA, timeout=120)
r.raise_for_status()
rows.append({"month": str(m), "count": r.json()["message"]["total-results"]})
pd.DataFrame(rows).to_csv(path, index=False)
time.sleep(0.2)
out = pd.DataFrame(rows)
return pd.Series(out["count"].values, index=pd.PeriodIndex(out["month"], freq="M"), name="journals").sort_index()
journals = crossref_monthly()
journals.groupby(journals.index.month).mean().round(-3).astype(int).rename("mean articles by calendar month").to_frame().T
```
One artefact matters later: records with only a publication year are counted on 1 January, so January is roughly double any other month. Comparisons involving Crossref therefore use February to August only.
### Step 3: the per-paper metadata
The official statistics say how many papers arrived. To say *where from* and *who from*, we need one row per paper. arXiv's own OAI-PMH interface serves that, but at roughly 1,300 records per two-minute request a full harvest of three million records would take days. Cornell publishes the same metadata as a regularly refreshed snapshot on Kaggle, which downloads in a few minutes.
```{python}
#| label: fetch-kaggle
KAGGLE_ZIP = RAW / "arxiv-metadata-kaggle.zip"
if not KAGGLE_ZIP.exists():
url = "https://www.kaggle.com/api/v1/datasets/download/Cornell-University/arxiv"
with requests.get(url, stream=True, timeout=600) as r:
r.raise_for_status()
with open(KAGGLE_ZIP, "wb") as fh:
for chunk in r.iter_content(1 << 22):
fh.write(chunk)
with zipfile.ZipFile(KAGGLE_ZIP) as z:
info = z.infolist()[0]
print(f"{info.filename}: {info.file_size/1e9:.1f} GB, snapshot written {info.date_time[0]}-{info.date_time[1]:02d}-{info.date_time[2]:02d}")
```
### Step 4: from 5.6 GB of JSON to a slim table
Each line of the snapshot is one paper. We keep only what the analysis needs and never unzip the file to disk: the records are streamed out of the archive, reduced to fifteen fields, and written as parquet parts of 400,000 papers each. Parts that already exist are skipped, so an interrupted run picks up where it stopped.
What each derived field means:
- `created` is the timestamp of version 1, the original submission. This matches how arXiv counts monthly submissions.
- `primary` is the first listed category; `n_cats` counts cross-lists.
- `submitter` is the display name of the account that uploaded the paper. It is a name, not an ID (see the limitations).
- `n_authors` comes from arXiv's own parsed author list.
- `accepted` flags comments that mention acceptance, publication or proceedings.
- `llm_marker` flags abstracts using any of a handful of words that LLMs of 2023 and 2024 were known to overuse.
```{python}
#| label: parse-fn
PAPERS = PROC / "papers.parquet"
COLS = ["id", "created", "n_versions", "primary", "n_cats", "submitter", "first_author", "n_authors",
"has_jref", "has_doi", "accepted", "n_pages", "abs_words", "llm_marker"]
MARKERS = re.compile(r"\b(delv(e|es|ed|ing)|underscor(e|es|ed|ing)|showcas(e|es|ed|ing)|pivotal|intricate|meticulous(ly)?)\b", re.I)
PAGES = re.compile(r"(\d{1,4})\s*pages?\b", re.I)
ACCEPTED = re.compile(r"accepted|to appear|published|proceedings|camera[- ]ready", re.I)
def parse_record(d):
cats = d["categories"].split()
authors = d["authors_parsed"]
comments = d["comments"] or ""
pages = PAGES.search(comments)
return (
d["id"], d["versions"][0]["created"], len(d["versions"]), cats[0], len(cats),
d["submitter"] or "", (authors[0][0] + ", " + authors[0][1]).strip(", ") if authors else "",
len(authors), bool(d["journal-ref"]), bool(d["doi"]), bool(ACCEPTED.search(comments)),
int(pages.group(1)) if pages else -1, len(d["abstract"].split()), bool(MARKERS.search(d["abstract"])),
)
def part_done(k):
"""A part counts as done if it exists and was written with the current set of columns."""
p = PARTS / f"part-{k:03d}.parquet"
return p.exists() and pq.read_schema(p).names == COLS
def parse_snapshot(zip_path=KAGGLE_ZIP, chunk=400_000, budget_s=None):
"""Stream the snapshot and write one parquet part per `chunk` records. Returns True when complete.
`budget_s` stops cleanly after that many seconds, for environments with short time limits."""
t0, rows, part = time.time(), [], 0
flush = lambda: pd.DataFrame(rows, columns=COLS).to_parquet(PARTS / f"part-{part:03d}.parquet", index=False)
with zipfile.ZipFile(zip_path) as z, z.open(z.infolist()[0]) as f:
have = part_done(part)
for i, line in enumerate(f):
if i and i % chunk == 0:
if not have:
flush()
rows, part = [], part + 1
have = part_done(part)
if budget_s and not have and time.time() - t0 > budget_s:
return False
if not have:
rows.append(parse_record(orjson.loads(line)))
if not have:
flush()
return True
```
The last piece of cleaning is the subject grouping. arXiv has eight top-level groups; physics is spread over a dozen archive prefixes (`astro-ph`, `hep-th`, `quant-ph` and so on) and a few pre-2000 archives were later folded into today's categories.
```{python}
#| label: load
PHYSICS = {"astro-ph", "cond-mat", "gr-qc", "hep-ex", "hep-lat", "hep-ph", "hep-th", "nlin",
"nucl-ex", "nucl-th", "physics", "quant-ph", "math-ph"}
LEGACY = {"cmp-lg": "cs", "alg-geom": "math", "dg-ga": "math", "funct-an": "math", "q-alg": "math"}
GROUPS = ["cs", "math", "physics", "stat", "eess", "econ", "q-bio", "q-fin"]
def group_of(cat):
archive = cat.split(".")[0]
if archive in PHYSICS:
return "physics"
return LEGACY.get(archive, archive if archive in GROUPS else "physics") # remaining legacy archives are physics
def build_papers():
df = pd.concat([pd.read_parquet(p) for p in sorted(PARTS.glob("part-*.parquet"))], ignore_index=True)
df["created"] = pd.to_datetime(df["created"], format="%a, %d %b %Y %H:%M:%S GMT")
df["group"] = df["primary"].map(group_of).astype("category")
df["primary"] = df["primary"].astype("category")
df["submitter"] = df["submitter"].astype("category")
df["month"] = df["created"].dt.to_period("M")
df.to_parquet(PAPERS, index=False)
return df
if not PAPERS.exists():
parse_snapshot()
build_papers()
papers = pd.read_parquet(PAPERS)
papers["year"] = papers["created"].dt.year
print(f"{len(papers):,} papers, first submitted {papers.created.min():%Y-%m-%d} to {papers.created.max():%Y-%m-%d}")
papers.sample(5, random_state=1)[["id", "created", "primary", "group", "n_authors", "n_versions", "accepted", "abs_words"]]
```
### Step 5: does the snapshot agree with arXiv's own counts?
Before trusting the per-paper table, compare its monthly counts with the official series.
```{python}
#| label: validate
snap = papers.groupby("month").size()
check = pd.DataFrame({"official": official, "snapshot": snap}).loc["2025-10":"2026-09"]
check["snapshot / official"] = (check["snapshot"] / check["official"]).round(3)
LAST = check.index[check["snapshot / official"] > 0.9][-1] # last month the snapshot fully covers
band = (snap / official).loc["2019-01":str(LAST)]
check
```
```{python}
#| label: window
M = LAST.month # compare the same months of every year: January to M
def ytd(df, col="created"):
return df[df[col].dt.month <= M]
W = ytd(papers[(papers.year >= 2015) & (papers.month <= LAST)]).copy()
Y1, Y0 = LAST.year, LAST.year - 1
n = W.groupby("year").size()
print(f"Comparison window: January to {LAST.strftime('%B')} of each year. {Y1}: {n[Y1]:,} papers; {Y0}: {n[Y0]:,}.")
```
Month by month the snapshot lands within `{python} f"{(band - 1).abs().max():.0%}"` of the official count (small differences come from papers announced or reclassified across a month boundary). The final month is incomplete because the snapshot was taken before it ended, so the per-paper analysis compares January to `{python} LAST.strftime('%B')` of each year: the same eight months, fully covered, with no seasonal distortion.
### Step 6: the derived tables
Everything in the findings below comes from the tables built here. First, growth by subject group and by category.
```{python}
#| label: derive-fields
def growth_table(counts, a, b):
"""counts: rows = year, columns = some split. Returns added papers, growth and share of total growth from a to b."""
t = pd.DataFrame({a: counts.loc[a], b: counts.loc[b]}).fillna(0).astype(int)
t["added"] = t[b] - t[a]
t["growth"] = t[b] / t[a] - 1
t["share of growth"] = t["added"] / t["added"].sum()
return t.sort_values("added", ascending=False)
by_group = W.groupby(["year", "group"], observed=True).size().unstack()
by_cat = W.groupby(["year", "primary"], observed=True).size().unstack(fill_value=0)
g_group = growth_table(by_group, Y0, Y1)
g_cat = growth_table(by_cat, Y0, Y1)
g_group.style.format({"growth": "{:+.1%}", "share of growth": "{:.1%}", Y0: "{:,}", Y1: "{:,}", "added": "{:+,}"})
```
Next, who is submitting. Three splits of the same papers: how long the submitting account has been on arXiv (years since its first paper anywhere in the snapshot), how many authors the paper has, and how many papers the account submitted inside the window.
```{python}
#| label: derive-submitters
first_year = papers.groupby("submitter", observed=True)["created"].transform("min").dt.year
S = W[W["submitter"] != ""].copy() # 0.5% of records have no submitter name
S["tenure"] = pd.cut(S["year"] - first_year.loc[S.index], [-1, 0, 4, 100], labels=["first year on arXiv", "1-4 years", "5+ years"])
S["team"] = pd.cut(S["n_authors"], [0, 1, 3, 6, 10**6], labels=["1 author", "2-3", "4-6", "7+"])
in_window = S.groupby(["year", "submitter"], observed=True)["id"].transform("size")
S["volume"] = pd.cut(in_window, [0, 1, 2, 4, 9, 10**6], labels=["1 paper", "2", "3-4", "5-9", "10+"])
g_tenure = growth_table(S.groupby(["year", "tenure"], observed=True).size().unstack(), Y0, Y1).sort_index()
g_team = growth_table(S.groupby(["year", "team"], observed=True).size().unstack(), Y0, Y1).sort_index()
g_volume = growth_table(S.groupby(["year", "volume"], observed=True).size().unstack(), Y0, Y1).sort_index()
per_year = S.groupby("year").agg(papers=("id", "size"), submitters=("submitter", "nunique"))
per_year["papers per submitter"] = per_year["papers"] / per_year["submitters"]
per_year["share from 5+ accounts"] = S[in_window >= 5].groupby("year").size() / per_year["papers"]
per_year["solo share"] = S[S.n_authors == 1].groupby("year").size() / per_year["papers"]
solo = S[S.n_authors == 1].groupby(["year", "submitter"], observed=True).size()
solo_prolific = pd.DataFrame({
"accounts with 5+ solo papers": (solo >= 5).groupby("year").sum(),
"accounts with 10+ solo papers": (solo >= 10).groupby("year").sum(),
"papers from 5+ solo accounts": solo[solo >= 5].groupby("year").sum(),
})
# arXiv's October 2026 rule: at most two new submissions per account per calendar month.
per_month = S.groupby(["year", "month", "submitter"], observed=True).size()
over_cap = (per_month - 2).clip(lower=0).groupby("year").sum() / per_year["papers"]
accounts_over = (per_month > 2).groupby("year").sum()
per_year.tail(3).round(3)
```
Then the monthly series used in the charts, and the comparison with other venues over February to `{python} LAST.strftime('%B')`.
```{python}
#| label: derive-series
P = papers[(papers.month >= "2015-01") & (papers.month <= LAST)]
monthly = P.groupby("month").agg(
papers=("id", "size"), solo=("n_authors", lambda s: (s == 1).mean()),
marker=("llm_marker", "mean"), accepted=("accepted", "mean"), abs_words=("abs_words", "median"))
official_full = official.iloc[:-1] # drop the partial current month
def feb_to_m(s):
s = s[(s.index.month >= 2) & (s.index.month <= M) & (s.index.year >= 2019)]
return s.groupby(s.index.year).sum()
venues = pd.DataFrame({"arXiv": feb_to_m(official_full), "bioRxiv": feb_to_m(biorxiv),
"medRxiv": feb_to_m(medrxiv), "Journal articles (Crossref)": feb_to_m(journals)})
venue_growth = pd.DataFrame({f"{Y0} to {Y1}": venues.loc[Y1] / venues.loc[Y0] - 1,
f"{Y1 - 3} to {Y1}": venues.loc[Y1] / venues.loc[Y1 - 3] - 1})
yoy = (n / n.shift(1) - 1)
K = dict(
n1=f"{n[Y1]:,}", n0=f"{n[Y0]:,}", yoy=f"{yoy[Y1]:.0%}", yoy_prev=f"{yoy[Y0]:.0%}",
sep=f"{official_full.iloc[-1]:,}", sep_month=official_full.index[-1].strftime("%B %Y"),
sep_x=f"{official_full.iloc[-1] / official_full.iloc[-25]:.1f}",
cs_share=f"{g_group.loc['cs', 'share of growth']:.0%}", cs_g=f"{g_group.loc['cs', 'growth']:.0%}",
math_g=f"{g_group.loc['math', 'growth']:.0%}", phys_g=f"{g_group.loc['physics', 'growth']:.0%}",
ai_x=f"{by_cat.loc[Y1, 'cs.AI'] / by_cat.loc[Y0, 'cs.AI']:.1f}",
debut_share=f"{g_tenure.loc['first year on arXiv', 'share of growth']:.0%}",
vet_share=f"{g_tenure.loc['5+ years', 'share of growth']:.0%}",
pps1=f"{per_year.loc[Y1, 'papers per submitter']:.2f}", pps0=f"{per_year.loc[Y0, 'papers per submitter']:.2f}",
solo_g=f"{g_team.loc['1 author', 'growth']:.0%}", solo_share=f"{g_team.loc['1 author', 'share of growth']:.0%}",
multi_g=f"{(g_team[Y1].sum() - g_team.loc['1 author', Y1]) / (g_team[Y0].sum() - g_team.loc['1 author', Y0]) - 1:.0%}",
sp1=f"{solo_prolific.loc[Y1, 'accounts with 5+ solo papers']:,}", sp0=f"{solo_prolific.loc[Y0, 'accounts with 5+ solo papers']:,}",
sp_papers=f"{solo_prolific.loc[Y1, 'papers from 5+ solo accounts']:,}",
sp_pct=f"{solo_prolific.loc[Y1, 'papers from 5+ solo accounts'] / per_year.loc[Y1, 'papers']:.1%}",
bio=f"{venue_growth.iloc[1, 0]:.0%}", med=f"{venue_growth.iloc[2, 0]:.0%}", jrn=f"{venue_growth.iloc[3, 0]:.0%}",
arx=f"{venue_growth.iloc[0, 0]:.0%}",
cap1=f"{over_cap[Y1]:.1%}", cap0=f"{over_cap[Y0]:.1%}",
marker_peak=f"{monthly.marker.max():.1%}", marker_peak_m=monthly.marker.idxmax().strftime("%B %Y"),
marker_now=f"{monthly.marker.iloc[-1]:.1%}",
)
venue_growth.style.format("{:+.1%}")
```
## The short answer
::: {.callout-note appearance="simple"}
**Both stories are partly right, and the metadata shows which part of each.**
1. **The lift-off is real, and it is steepest on arXiv.** January to `{python} LAST.strftime('%B %Y')` brought `{python} K['n1']` new papers, `{python} K['yoy']` more than the same months a year earlier. Over February to August, bioRxiv grew `{python} K['bio']`, medRxiv `{python} K['med']`, and journal articles registered with Crossref `{python} K['jrn']`.
2. **It is not only an AI-papers story.** Computer science supplied `{python} K['cs_share']` of the added papers and cs.AI alone grew `{python} K['ai_x']`x, but mathematics grew `{python} K['math_g']` and physics `{python} K['phys_g']` after years of single-digit growth.
3. **It is mostly established accounts submitting more, not a flood of newcomers.** Only `{python} K['debut_share']` of the growth came from accounts in their first year on arXiv; `{python} K['vet_share']` came from accounts at least five years old. Papers per account, flat for a decade, jumped from `{python} K['pps0']` to `{python} K['pps1']`.
4. **The mill-like signature exists, and it is small.** Single-author papers rose `{python} K['solo_g']` (everything else: `{python} K['multi_g']`), reversing a decade-long decline. Accounts posting five or more solo papers in eight months went from `{python} K['sp0']` to `{python} K['sp1']`. Their solo papers number `{python} K['sp_papers']`, or `{python} K['sp_pct']` of the total.
5. **Metadata cannot grade quality.** It shows output accelerating much faster than peer-reviewed venues are absorbing it. Whether that output is good research is a question for outcomes we can only observe later.
:::
## The lift-off is real
```{python}
#| label: fig-liftoff
#| fig-cap: "New arXiv submissions per month (official arXiv statistics)."
s = official_full.loc["2008-01":]
x = s.index.to_timestamp()
fig, ax = plt.subplots()
ax.plot(x, s.values, color="#c3c2b7", lw=1, label="Monthly")
ax.plot(x, s.rolling(12).mean().values, color=BLUE, label="12-month average")
ax.scatter(x[-1], s.iloc[-1], s=36, color=BLUE, zorder=3, edgecolor="#fcfcfb", linewidth=2)
ax.annotate(f"{s.index[-1].strftime('%b %Y')}: {s.iloc[-1]:,}", (x[-1], s.iloc[-1]), xytext=(-8, 0),
textcoords="offset points", ha="right", va="center", color=INK)
ax.yaxis.set_major_formatter(thousands); ax.set_ylim(0); ax.legend(loc="upper left")
ax.set_title("arXiv submissions per month")
plt.show()
```
Through the 2010s arXiv grew by 5 to 14 percent a year, and it nearly stalled in 2021 and 2022. The first eight months of `{python} Y1` grew `{python} K['yoy']`, after `{python} K['yoy_prev']` the year before. `{python} K['sep_month']`, the latest complete month in arXiv's own statistics, brought `{python} K['sep']` submissions, `{python} K['sep_x']` times the figure of two years earlier. The "lift-off" in the LinkedIn post is not a plotting artefact.
::: {.thorough}
Year-over-year growth for the same January to `{python} LAST.strftime('%B')` window, from the per-paper table:
```{python}
#| label: yoy-table
pd.DataFrame({"papers": n, "growth": yoy}).loc[2016:].style.format({"papers": "{:,}", "growth": "{:+.1%}"})
```
:::
## Where the growth comes from
```{python}
#| label: fig-fields
#| fig-cap: "Papers added between the two January to August windows, by subject group and by primary category."
fig, axes = plt.subplots(1, 2, figsize=(9.2, 4.2), gridspec_kw={"wspace": 0.55})
names = {"cs": "Computer science", "physics": "Physics", "math": "Mathematics", "stat": "Statistics",
"econ": "Economics", "eess": "Electrical eng.", "q-fin": "Quant. finance", "q-bio": "Quant. biology"}
for ax, t, title, lab in [(axes[0], g_group, "By subject group", lambda i: names[i]),
(axes[1], g_cat.head(10), "Top 10 categories", str)]:
t = t.iloc[::-1]
ax.barh([lab(i) for i in t.index], t["added"], color=BLUE, height=0.62)
for y, (a, g) in enumerate(zip(t["added"], t["growth"])):
ax.text(max(a, 0), y, f" {a:+,} ({g:+.0%})", va="center", ha="left", color=INK2, fontsize=9)
ax.set_title(title); ax.grid(False); ax.set_xticks([]); ax.spines["bottom"].set_visible(False)
ax.set_xlim(min(0, t["added"].min()), t["added"].max() * 1.75); ax.tick_params(length=0)
fig.suptitle(f"Papers added, Jan to {LAST.strftime('%b')} {Y0} vs {Y1} (growth in brackets)", x=0.02, ha="left", fontsize=10, color=INK2)
plt.show()
```
Computer science accounts for `{python} K['cs_share']` of the added papers, and one category stands out: cs.AI grew `{python} K['ai_x']`x in a single year. arXiv singled out the same category when it [introduced submission rate limits on 1 October 2026](https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/).
The more surprising result is outside CS. Mathematics grew `{python} K['math_g']` and physics `{python} K['phys_g']`, in fields that had been growing by a few percent a year. Combinatorics (math.CO) grew faster than computer vision. A surge that reaches pure mathematics is not explained by "more people are working on AI". Something changed in how quickly papers get written across fields.
::: {.thorough}
The full group table is in Step 6. The twenty categories that added the most papers:
```{python}
#| label: categories-table
g_cat.head(20).style.format({"growth": "{:+.1%}", "share of growth": "{:.1%}", Y0: "{:,}", Y1: "{:,}", "added": "{:+,}"})
```
:::
## Who is submitting more
```{python}
#| label: fig-who
#| fig-cap: "Left: papers per submitting account within each January to August window. Right: which accounts supplied the growth."
fig, axes = plt.subplots(1, 2, figsize=(9.2, 3.8), gridspec_kw={"wspace": 0.3})
ax = axes[0]
pps = per_year["papers per submitter"]
ax.plot(pps.index, pps.values, color=BLUE, marker="o", ms=5, markeredgecolor="#fcfcfb")
ax.annotate(f"{pps.iloc[-1]:.2f}", (pps.index[-1], pps.iloc[-1]), xytext=(-6, 4), textcoords="offset points", ha="right")
ax.set_ylim(1.3, 1.65); ax.set_title("Papers per account"); ax.set_xticks(range(2015, Y1 + 1, 2))
ax = axes[1]
ax.bar(g_tenure.index.astype(str), g_tenure["share of growth"], color=BLUE, width=0.55)
for i, v in enumerate(g_tenure["share of growth"]):
ax.text(i, v + 0.01, f"{v:.0%}", ha="center", color=INK)
ax.yaxis.set_major_formatter(pct0); ax.set_ylim(0, 0.6); ax.set_title(f"Share of {Y0} to {Y1} growth, by account age")
plt.show()
```
If paper mills or first-time hobbyists were driving the surge, new accounts would dominate the growth. They do not. Accounts in their first year supplied `{python} K['debut_share']` of the added papers, and their output grew more slowly than anyone else's, which fits arXiv's [tightened endorsement rules for new submitters](https://blog.arxiv.org/2026/01/21/attention-authors-updated-endorsement-policy/) in January 2026. Half of the growth (`{python} K['vet_share']`) came from accounts that have been on arXiv for five years or more.
What changed is the pace. For ten years the average account submitted about 1.45 papers in an eight-month window, almost without variation. In `{python} Y1` that rose to `{python} K['pps1']`. The people who were already writing papers are writing more of them.
::: {.thorough}
The three splits behind this section. Accounts that submitted five or more papers in the window are where output grew fastest:
```{python}
#| label: who-table
fmt = {"growth": "{:+.1%}", "share of growth": "{:.1%}", Y0: "{:,}", Y1: "{:,}", "added": "{:+,}"}
pd.concat({"Account age": g_tenure, "Authors on the paper": g_team, "Papers by the account in the window": g_volume}).style.format(fmt)
```
:::
## The solo-author break
```{python}
#| label: fig-solo
#| fig-cap: "Left: share of new papers with a single author, by month. Right: accounts that submitted five or more single-author papers in a January to August window."
fig, axes = plt.subplots(1, 2, figsize=(9.2, 3.8), gridspec_kw={"wspace": 0.3})
ax = axes[0]
sm = monthly["solo"].rolling(3).mean()
ax.plot(sm.index.to_timestamp(), sm.values, color=BLUE)
ax.yaxis.set_major_formatter(pct0); ax.set_ylim(0, 0.25); ax.set_title("Single-author share (3-month average)")
ax = axes[1]
sp = solo_prolific["accounts with 5+ solo papers"]
ax.bar(sp.index, sp.values, color=BLUE, width=0.62)
ax.text(sp.index[-1], sp.iloc[-1] + 8, f"{sp.iloc[-1]:,}", ha="center"); ax.text(sp.index[-2], sp.iloc[-2] + 8, f"{sp.iloc[-2]:,}", ha="center")
ax.set_title("Accounts with 5+ solo papers"); ax.set_xticks(range(2015, Y1 + 1, 2))
plt.show()
```
This is the clearest structural break in the data. Research had been getting steadily more collaborative: the single-author share of new papers fell from about 20 percent in 2015 to under 11 percent in 2025. In `{python} Y1` it turned around. Single-author papers grew `{python} K['solo_g']` while multi-author papers grew `{python} K['multi_g']`, and solo papers alone account for `{python} K['solo_share']` of all the growth.
At the extreme end is the pattern Srijit described. The number of accounts posting five or more single-author papers in eight months sat around 100 for a decade, reached `{python} K['sp0']` in `{python} Y0`, and then `{python} K['sp1']` in `{python} Y1`. One person producing a paper every few weeks, alone, was rare enough to be a curiosity. It no longer is.
It is also a small part of the whole. Their solo papers number `{python} K['sp_papers']`, or `{python} K['sp_pct']` of the window. arXiv's new cap of two submissions per account per month would have held back `{python} K['cap1']` of `{python} Y1`'s papers (it was `{python} K['cap0']` the year before). The high-frequency tail is real and growing fast, but removing it entirely would leave nearly all of the surge in place.
::: {.thorough}
```{python}
#| label: solo-table
per_year.join(solo_prolific).join(over_cap.rename("share above 2/month")).style.format({
"papers": "{:,}", "submitters": "{:,}", "papers per submitter": "{:.3f}", "share from 5+ accounts": "{:.1%}",
"solo share": "{:.1%}", "papers from 5+ solo accounts": "{:,}", "share above 2/month": "{:.1%}"})
```
No accounts are named here on purpose. A submitter name is not proof of anything: prolific legitimate accounts exist (proceedings series, large collaborations), and common names merge several people.
:::
## Are credible venues growing the same way?
```{python}
#| label: fig-venues
#| fig-cap: "Output over February to August of each year, indexed to 2023. Crossref counts are journal articles by publication date."
base = Y1 - 3
idx = venues.loc[base:] / venues.loc[base] * 100
fig, ax = plt.subplots(figsize=(7.6, 4.0))
for col, c in zip(idx.columns, [BLUE, ORANGE, AQUA, YELLOW]):
ax.plot(idx.index, idx[col], color=c, marker="o", ms=5, markeredgecolor="#fcfcfb", label=col)
ax.annotate(f"{col} {idx[col].iloc[-1]:.0f}", (Y1, idx[col].iloc[-1]), xytext=(8, 0), textcoords="offset points", va="center", color=INK)
ax.set_xticks(idx.index); ax.set_xlim(base - 0.15, Y1 + 1.6); ax.set_ylim(95, None)
ax.legend(loc="upper left"); ax.set_title(f"Growth since {base} ({base} = 100)")
plt.show()
```
This was the test proposed in the comments, and the answer is no. Over three years arXiv grew about `{python} f"{venue_growth.iloc[0, 1]:.0%}"`, while journal articles registered with Crossref grew `{python} f"{venue_growth.iloc[3, 1]:.0%}"`. In the latest year arXiv grew `{python} K['arx']` against `{python} K['jrn']` for journals. The life-science preprint servers sit in between. medRxiv comes closest, at `{python} K['med']` in the latest year, but it is a twentieth of arXiv's size and was already growing almost that fast; neither server shows a jump in `{python} Y1` like arXiv's.
Two cautions. Crossref counts every journal with a DOI, so "credible, defined broadly" is doing a lot of work, and journal growth includes paper-mill output of its own. And journals lag preprints: a paper posted in 2026 may be published in 2027. A gap this wide still means that preprints are being produced far faster than the peer-review system is currently absorbing them. Either a wave of publications is on its way, or a growing share of arXiv will never be reviewed by anyone.
::: {.thorough}
```{python}
#| label: venues-table
venues.style.format("{:,}")
```
:::
## What the abstracts can and cannot tell us
```{python}
#| label: fig-marker
#| fig-cap: "Share of new abstracts using at least one of: delve, underscore, showcase, pivotal, intricate, meticulous (any inflection)."
mk = monthly["marker"].loc["2019-01":]
fig, ax = plt.subplots()
ax.plot(mk.index.to_timestamp(), mk.values, color=BLUE)
ax.axvline(pd.Timestamp("2022-11-30"), color=MUTED, lw=1, ls=(0, (3, 3)))
ax.text(pd.Timestamp("2022-11-01"), mk.max(), "ChatGPT released ", ha="right", va="top", color=INK2)
ax.yaxis.set_major_formatter(pct0); ax.set_ylim(0); ax.set_title("Abstracts with LLM marker words")
plt.show()
```
An obvious idea is to detect AI-written text directly. Researchers have done this through "excess vocabulary": words such as *delve* that became abruptly more common after ChatGPT appeared ([Kobak et al.](https://arxiv.org/abs/2406.07016)). The same handful of words in arXiv abstracts shows the rise clearly, peaking at `{python} K['marker_peak']` of abstracts in `{python} K['marker_peak_m']`.
Then it falls away, back to `{python} K['marker_now']` by `{python} LAST.strftime('%B %Y')`, during the exact period when submissions took off. Nobody believes LLM use declined in 2026. Models changed their habits and authors learned to edit out the tells. A fixed word list is a detector for one generation of models, and it says nothing about quality in any case: a careful paper polished by an LLM and a fabricated one both trip it.
Abstracts did get longer, though: the median went from `{python} f"{monthly.abs_words.loc[f'{Y0}-01':f'{Y0}-08'].median():.0f}"` words in early `{python} Y0` to `{python} f"{monthly.abs_words.iloc[-3:].median():.0f}"` by mid `{python} Y1`, after creeping up by a word or two a year.
::: {.thorough}
### Why publication status cannot settle it yet
The natural quality measure is whether a preprint ends up published. The snapshot has three signals: a journal reference, a DOI, and comments mentioning acceptance. All three are added *after* the fact, so recent cohorts always look worse than old ones regardless of quality:
```{python}
#| label: censoring-table
c = W.groupby("year").agg(journal_ref=("has_jref", "mean"), doi=("has_doi", "mean"), accepted_in_comments=("accepted", "mean"))
c.loc[2018:].style.format("{:.1%}")
```
The decline starts years before the surge and steepens toward the present, which is what right-censoring looks like. Comparing cohorts fairly needs each one measured at the same age, which means keeping successive snapshots. That is the follow-up worth doing: re-run this notebook in a year and compare the `{python} Y1` cohort's publication rate at twelve months with the `{python} Y0` cohort's today.
:::
## So: acceleration or slop?
The data rule out the simple version of each story.
It is not a flood of outsiders and mills. The growth comes mainly from people with a track record on arXiv, across computer science, mathematics and physics at once, and the accounts that behave like mills are a few percent of the total at most.
It is also not business as usual at higher speed. Papers per account jumped after a decade of stability, solo work reversed a long decline, and peer-reviewed output did not move with it. The same researchers are producing papers faster than before, increasingly alone, and faster than journals and conferences are certifying them.
Whether that is acceleration of *research* or only of *papers* is not something titles, author lists and abstracts can decide. arXiv's moderators, who read the submissions, describe a rise in ["thin papers of narrow scope"](https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/) and work split into smaller pieces, and have responded with endorsement rules and rate limits. The measurements that would settle it are outcomes: what share of the `{python} Y1` cohort is published a year on, how it is cited, and how often it is withdrawn. Those can be measured, just not yet.
## Limitations {.thorough}
- **Submitter names are not identities.** The snapshot has display names, not account IDs. Common names merge several people and inflate "papers per account"; a person who changes their display name is split. This affects every year, but collisions grow with volume, so the level of the prolific-account counts is less reliable than their direction. The solo-author share does not depend on names at all.
- **Account age is measured from the first paper under that name** in the snapshot, with the same caveat.
- **The comparison window is January to `{python} LAST.strftime('%B')`.** `{python} K['sep_month']` appears only in the official monthly count; the per-paper snapshot was taken before that month ended.
- **Crossref dates are publication dates as deposited by publishers.** Year-only dates land on 1 January (excluded here), and recent months can still gain records as publishers deposit late, which would slightly understate journal growth.
- **Primary category only.** Cross-listed papers are counted once, under their first category.
- **No full text and no moderation data.** Rejected submissions never appear in the snapshot, so this is the growth that survived arXiv's moderation.
## Reproducing this {.thorough}
```bash
conda env create -f environment.yml && conda activate arxiv-analysis
quarto render arxiv-liftoff.qmd
```
The first render downloads about 1.8 GB and takes several minutes to parse; later renders reuse `data/` and finish in seconds. Delete `data/raw/` to refresh every source.