arXiv has reached lift-off. Is it acceleration or slop?

Where the surge in preprints is coming from, and what the metadata can and cannot tell us

arXiv
AI
Data Analysis
Author

Trey Saddler

Published

October 1, 2026

A post on LinkedIn recently put the question plainly: monthly submissions to arXiv have “reached lift-off”, and either AI has meaningfully accelerated research, or preprints are filling up with low-effort slop. How would we measure which is true? In the comments, Srijit Seal suggested a place to start: work out where the growth is coming from and how fast it is expanding, and check whether credible venues show the same explosion. If they do not, look for mills that “just submit papers every hour”.

This article does that analysis with public data. Use the switch under the title to move between the takeaways and the full walkthrough, which shows how each dataset was fetched, every processing step, and the code behind every number.

The analysis is a single Quarto document running on the Jupyter kernel. This first cell sets up paths and one shared chart style.

Show the code
from pathlib import Path
import json, re, time, zipfile

import matplotlib.pyplot as plt
import matplotlib.ticker as mtick
import numpy as np
import orjson
import pandas as pd
import pyarrow.parquet as pq
import requests

DATA = Path("arxiv-liftoff/data")  # relative to posts/; not committed to the site repo
RAW, PROC = DATA / "raw", DATA / "processed"
PARTS = PROC / "parts"
for d in (RAW, PROC, PARTS):
    d.mkdir(parents=True, exist_ok=True)

UA = {"User-Agent": "arxiv-liftoff-analysis (mailto:treyosaddler@gmail.com)"}

# One visual system for every chart: thin marks, recessive axes, a fixed colour order.
BLUE, ORANGE, AQUA, YELLOW = "#2a78d6", "#eb6834", "#1baf7a", "#eda100"
INK, INK2, MUTED, GRID = "#0b0b0b", "#52514e", "#898781", "#e1e0d9"
plt.rcParams.update({
    "figure.figsize": (7.6, 3.9), "figure.facecolor": "#fcfcfb", "axes.facecolor": "#fcfcfb",
    "axes.spines.top": False, "axes.spines.right": False, "axes.spines.left": False,
    "axes.edgecolor": "#c3c2b7", "axes.grid": True, "axes.grid.axis": "y", "grid.color": GRID,
    "grid.linewidth": 0.8, "axes.axisbelow": True, "xtick.color": MUTED, "ytick.color": MUTED,
    "xtick.labelcolor": INK2, "ytick.labelcolor": INK2, "ytick.left": False, "font.size": 10,
    "axes.titlesize": 11, "axes.titleweight": "bold", "axes.titlelocation": "left",
    "axes.labelcolor": INK2, "text.color": INK, "lines.linewidth": 2, "legend.frameon": False,
    "svg.fonttype": "none",
})
thousands = mtick.FuncFormatter(lambda v, _: f"{v:,.0f}")
pct0 = mtick.PercentFormatter(1, decimals=0)

The data

Four public sources, each cached under data/raw/ the first time it is fetched so that re-rendering does not hit anyone’s servers again.

Source What it gives us How it is fetched
arXiv monthly submission statistics Official count of new submissions per month since 1991 One CSV download
arXiv metadata snapshot on Kaggle One JSON record per paper: submitter, authors, categories, versions, comments, abstract One 1.8 GB zip (5.6 GB of JSON lines)
bioRxiv and medRxiv summary APIs New preprints per month on the two largest life-science servers One JSON call each
Crossref REST API Number of journal articles registered, by publication month One count query per month

Step 1: the small sources

Show the code
def fetch(url, name):
    """Download `url` to data/raw/`name` once; later runs reuse the cached copy."""
    path = RAW / name
    if not path.exists():
        r = requests.get(url, headers=UA, timeout=120)
        r.raise_for_status()
        path.write_bytes(r.content)
    return path

official = pd.read_csv(fetch("https://arxiv.org/stats/get_monthly_submissions", "arxiv_monthly_submissions.csv"))
official["month"] = pd.PeriodIndex(official["month"], freq="M")
official = official.set_index("month")["submissions"]

def rxiv_monthly(server):
    j = json.loads(fetch(f"https://api.{server}.org/sum/m", f"{server}_monthly.json").read_text())
    key = next(k for k in j if "statistics" in k)
    d = pd.DataFrame(j[key])
    return pd.Series(d["new_papers"].astype(int).values, index=pd.PeriodIndex(d["month"], freq="M"), name=server)

biorxiv, medrxiv = rxiv_monthly("biorxiv"), rxiv_monthly("medrxiv")
official.tail(4).to_frame()
submissions
month
2026-07 29687
2026-08 31173
2026-09 40363
2026-10 2210

The official arXiv series ends with a partial month (the file was pulled on 1 October 2026), so everything below stops at the last complete month.

Step 2: journal output from Crossref

Crossref is the DOI registry most journals use. Asking for zero rows with a filter returns just the count of matching records, so one polite request per month gives a monthly series of journal articles by publication date. The loop is resumable: each answer is written to the CSV straight away.

Show the code
def crossref_monthly(kind="journal-article", start="2015-01", end="2026-09"):
    path = RAW / f"crossref_{kind}_monthly.csv"
    have = pd.read_csv(path) if path.exists() else pd.DataFrame(columns=["month", "count"])
    rows, seen = have.to_dict("records"), set(have["month"])
    for m in pd.period_range(start, end, freq="M"):
        if str(m) in seen:
            continue
        flt = f"type:{kind},from-pub-date:{m.start_time:%Y-%m-%d},until-pub-date:{m.end_time:%Y-%m-%d}"
        r = requests.get("https://api.crossref.org/works", params={"rows": 0, "filter": flt}, headers=UA, timeout=120)
        r.raise_for_status()
        rows.append({"month": str(m), "count": r.json()["message"]["total-results"]})
        pd.DataFrame(rows).to_csv(path, index=False)
        time.sleep(0.2)
    out = pd.DataFrame(rows)
    return pd.Series(out["count"].values, index=pd.PeriodIndex(out["month"], freq="M"), name="journals").sort_index()

journals = crossref_monthly()
journals.groupby(journals.index.month).mean().round(-3).astype(int).rename("mean articles by calendar month").to_frame().T
month 1 2 3 4 5 6 7 8 9 10 11 12
mean articles by calendar month 922000 301000 365000 355000 346000 413000 356000 324000 376000 367000 350000 473000

One artefact matters later: records with only a publication year are counted on 1 January, so January is roughly double any other month. Comparisons involving Crossref therefore use February to August only.

Step 3: the per-paper metadata

The official statistics say how many papers arrived. To say where from and who from, we need one row per paper. arXiv’s own OAI-PMH interface serves that, but at roughly 1,300 records per two-minute request a full harvest of three million records would take days. Cornell publishes the same metadata as a regularly refreshed snapshot on Kaggle, which downloads in a few minutes.

Show the code
KAGGLE_ZIP = RAW / "arxiv-metadata-kaggle.zip"
if not KAGGLE_ZIP.exists():
    url = "https://www.kaggle.com/api/v1/datasets/download/Cornell-University/arxiv"
    with requests.get(url, stream=True, timeout=600) as r:
        r.raise_for_status()
        with open(KAGGLE_ZIP, "wb") as fh:
            for chunk in r.iter_content(1 << 22):
                fh.write(chunk)

with zipfile.ZipFile(KAGGLE_ZIP) as z:
    info = z.infolist()[0]
print(f"{info.filename}: {info.file_size/1e9:.1f} GB, snapshot written {info.date_time[0]}-{info.date_time[1]:02d}-{info.date_time[2]:02d}")
arxiv-metadata-oai-snapshot.json: 5.6 GB, snapshot written 2026-09-26

Step 4: from 5.6 GB of JSON to a slim table

Each line of the snapshot is one paper. We keep only what the analysis needs and never unzip the file to disk: the records are streamed out of the archive, reduced to fifteen fields, and written as parquet parts of 400,000 papers each. Parts that already exist are skipped, so an interrupted run picks up where it stopped.

What each derived field means:

  • created is the timestamp of version 1, the original submission. This matches how arXiv counts monthly submissions.
  • primary is the first listed category; n_cats counts cross-lists.
  • submitter is the display name of the account that uploaded the paper. It is a name, not an ID (see the limitations).
  • n_authors comes from arXiv’s own parsed author list.
  • accepted flags comments that mention acceptance, publication or proceedings.
  • llm_marker flags abstracts using any of a handful of words that LLMs of 2023 and 2024 were known to overuse.
Show the code
PAPERS = PROC / "papers.parquet"
COLS = ["id", "created", "n_versions", "primary", "n_cats", "submitter", "first_author", "n_authors",
        "has_jref", "has_doi", "accepted", "n_pages", "abs_words", "llm_marker"]

MARKERS = re.compile(r"\b(delv(e|es|ed|ing)|underscor(e|es|ed|ing)|showcas(e|es|ed|ing)|pivotal|intricate|meticulous(ly)?)\b", re.I)
PAGES = re.compile(r"(\d{1,4})\s*pages?\b", re.I)
ACCEPTED = re.compile(r"accepted|to appear|published|proceedings|camera[- ]ready", re.I)

def parse_record(d):
    cats = d["categories"].split()
    authors = d["authors_parsed"]
    comments = d["comments"] or ""
    pages = PAGES.search(comments)
    return (
        d["id"], d["versions"][0]["created"], len(d["versions"]), cats[0], len(cats),
        d["submitter"] or "", (authors[0][0] + ", " + authors[0][1]).strip(", ") if authors else "",
        len(authors), bool(d["journal-ref"]), bool(d["doi"]), bool(ACCEPTED.search(comments)),
        int(pages.group(1)) if pages else -1, len(d["abstract"].split()), bool(MARKERS.search(d["abstract"])),
    )

def part_done(k):
    """A part counts as done if it exists and was written with the current set of columns."""
    p = PARTS / f"part-{k:03d}.parquet"
    return p.exists() and pq.read_schema(p).names == COLS

def parse_snapshot(zip_path=KAGGLE_ZIP, chunk=400_000, budget_s=None):
    """Stream the snapshot and write one parquet part per `chunk` records. Returns True when complete.
    `budget_s` stops cleanly after that many seconds, for environments with short time limits."""
    t0, rows, part = time.time(), [], 0
    flush = lambda: pd.DataFrame(rows, columns=COLS).to_parquet(PARTS / f"part-{part:03d}.parquet", index=False)
    with zipfile.ZipFile(zip_path) as z, z.open(z.infolist()[0]) as f:
        have = part_done(part)
        for i, line in enumerate(f):
            if i and i % chunk == 0:
                if not have:
                    flush()
                rows, part = [], part + 1
                have = part_done(part)
                if budget_s and not have and time.time() - t0 > budget_s:
                    return False
            if not have:
                rows.append(parse_record(orjson.loads(line)))
        if not have:
            flush()
    return True

The last piece of cleaning is the subject grouping. arXiv has eight top-level groups; physics is spread over a dozen archive prefixes (astro-ph, hep-th, quant-ph and so on) and a few pre-2000 archives were later folded into today’s categories.

Show the code
PHYSICS = {"astro-ph", "cond-mat", "gr-qc", "hep-ex", "hep-lat", "hep-ph", "hep-th", "nlin",
           "nucl-ex", "nucl-th", "physics", "quant-ph", "math-ph"}
LEGACY = {"cmp-lg": "cs", "alg-geom": "math", "dg-ga": "math", "funct-an": "math", "q-alg": "math"}
GROUPS = ["cs", "math", "physics", "stat", "eess", "econ", "q-bio", "q-fin"]

def group_of(cat):
    archive = cat.split(".")[0]
    if archive in PHYSICS:
        return "physics"
    return LEGACY.get(archive, archive if archive in GROUPS else "physics")  # remaining legacy archives are physics

def build_papers():
    df = pd.concat([pd.read_parquet(p) for p in sorted(PARTS.glob("part-*.parquet"))], ignore_index=True)
    df["created"] = pd.to_datetime(df["created"], format="%a, %d %b %Y %H:%M:%S GMT")
    df["group"] = df["primary"].map(group_of).astype("category")
    df["primary"] = df["primary"].astype("category")
    df["submitter"] = df["submitter"].astype("category")
    df["month"] = df["created"].dt.to_period("M")
    df.to_parquet(PAPERS, index=False)
    return df

if not PAPERS.exists():
    parse_snapshot()
    build_papers()
papers = pd.read_parquet(PAPERS)
papers["year"] = papers["created"].dt.year
print(f"{len(papers):,} papers, first submitted {papers.created.min():%Y-%m-%d} to {papers.created.max():%Y-%m-%d}")
papers.sample(5, random_state=1)[["id", "created", "primary", "group", "n_authors", "n_versions", "accepted", "abs_words"]]
3,182,867 papers, first submitted 1986-04-25 to 2026-09-24
id created primary group n_authors n_versions accepted abs_words
2898801 cond-mat/0603357 2006-03-13 16:37:57 cond-mat.mes-hall physics 3 1 False 87
2452679 2510.20853 2025-10-22 05:11:02 eess.AS eess 14 2 True 170
595069 1502.00409 2015-02-02 09:06:03 math.CO math 2 1 False 97
1245092 2002.07669 2020-02-18 15:59:17 astro-ph.SR physics 3 1 True 192
2333519 2505.22958 2025-05-29 00:51:18 math.AT math 1 1 False 78

Step 5: does the snapshot agree with arXiv’s own counts?

Before trusting the per-paper table, compare its monthly counts with the official series.

Show the code
snap = papers.groupby("month").size()
check = pd.DataFrame({"official": official, "snapshot": snap}).loc["2025-10":"2026-09"]
check["snapshot / official"] = (check["snapshot"] / check["official"]).round(3)
LAST = check.index[check["snapshot / official"] > 0.9][-1]     # last month the snapshot fully covers
band = (snap / official).loc["2019-01":str(LAST)]
check
official snapshot snapshot / official
month
2025-10 27692.0 27477.0 0.992
2025-11 23478.0 24398.0 1.039
2025-12 25075.0 24160.0 0.964
2026-01 23286.0 24135.0 1.036
2026-02 24290.0 24378.0 1.004
2026-03 30045.0 29578.0 0.984
2026-04 28197.0 28202.0 1.000
2026-05 31604.0 33047.0 1.046
2026-06 32040.0 31009.0 0.968
2026-07 29687.0 29995.0 1.010
2026-08 31173.0 31507.0 1.011
2026-09 40363.0 28044.0 0.695
Show the code
M = LAST.month                      # compare the same months of every year: January to M
def ytd(df, col="created"):
    return df[df[col].dt.month <= M]
W = ytd(papers[(papers.year >= 2015) & (papers.month <= LAST)]).copy()
Y1, Y0 = LAST.year, LAST.year - 1
n = W.groupby("year").size()
print(f"Comparison window: January to {LAST.strftime('%B')} of each year. {Y1}: {n[Y1]:,} papers; {Y0}: {n[Y0]:,}.")
Comparison window: January to August of each year. 2026: 231,851 papers; 2025: 181,969.

Month by month the snapshot lands within 5% of the official count (small differences come from papers announced or reclassified across a month boundary). The final month is incomplete because the snapshot was taken before it ended, so the per-paper analysis compares January to August of each year: the same eight months, fully covered, with no seasonal distortion.

Step 6: the derived tables

Everything in the findings below comes from the tables built here. First, growth by subject group and by category.

Show the code
def growth_table(counts, a, b):
    """counts: rows = year, columns = some split. Returns added papers, growth and share of total growth from a to b."""
    t = pd.DataFrame({a: counts.loc[a], b: counts.loc[b]}).fillna(0).astype(int)
    t["added"] = t[b] - t[a]
    t["growth"] = t[b] / t[a] - 1
    t["share of growth"] = t["added"] / t["added"].sum()
    return t.sort_values("added", ascending=False)

by_group = W.groupby(["year", "group"], observed=True).size().unstack()
by_cat = W.groupby(["year", "primary"], observed=True).size().unstack(fill_value=0)
g_group = growth_table(by_group, Y0, Y1)
g_cat = growth_table(by_cat, Y0, Y1)
g_group.style.format({"growth": "{:+.1%}", "share of growth": "{:.1%}", Y0: "{:,}", Y1: "{:,}", "added": "{:+,}"})
  2025 2026 added growth share of growth
group          
cs 83,191 111,151 +27,960 +33.6% 56.1%
physics 54,766 65,077 +10,311 +18.8% 20.7%
math 27,102 36,452 +9,350 +34.5% 18.7%
stat 4,426 5,936 +1,510 +34.1% 3.0%
econ 1,427 1,776 +349 +24.5% 0.7%
eess 8,309 8,571 +262 +3.2% 0.5%
q-fin 722 981 +259 +35.9% 0.5%
q-bio 2,026 1,907 -119 -5.9% -0.2%

Next, who is submitting. Three splits of the same papers: how long the submitting account has been on arXiv (years since its first paper anywhere in the snapshot), how many authors the paper has, and how many papers the account submitted inside the window.

Show the code
first_year = papers.groupby("submitter", observed=True)["created"].transform("min").dt.year
S = W[W["submitter"] != ""].copy()                       # 0.5% of records have no submitter name
S["tenure"] = pd.cut(S["year"] - first_year.loc[S.index], [-1, 0, 4, 100], labels=["first year on arXiv", "1-4 years", "5+ years"])
S["team"] = pd.cut(S["n_authors"], [0, 1, 3, 6, 10**6], labels=["1 author", "2-3", "4-6", "7+"])
in_window = S.groupby(["year", "submitter"], observed=True)["id"].transform("size")
S["volume"] = pd.cut(in_window, [0, 1, 2, 4, 9, 10**6], labels=["1 paper", "2", "3-4", "5-9", "10+"])

g_tenure = growth_table(S.groupby(["year", "tenure"], observed=True).size().unstack(), Y0, Y1).sort_index()
g_team = growth_table(S.groupby(["year", "team"], observed=True).size().unstack(), Y0, Y1).sort_index()
g_volume = growth_table(S.groupby(["year", "volume"], observed=True).size().unstack(), Y0, Y1).sort_index()

per_year = S.groupby("year").agg(papers=("id", "size"), submitters=("submitter", "nunique"))
per_year["papers per submitter"] = per_year["papers"] / per_year["submitters"]
per_year["share from 5+ accounts"] = S[in_window >= 5].groupby("year").size() / per_year["papers"]
per_year["solo share"] = S[S.n_authors == 1].groupby("year").size() / per_year["papers"]

solo = S[S.n_authors == 1].groupby(["year", "submitter"], observed=True).size()
solo_prolific = pd.DataFrame({
    "accounts with 5+ solo papers": (solo >= 5).groupby("year").sum(),
    "accounts with 10+ solo papers": (solo >= 10).groupby("year").sum(),
    "papers from 5+ solo accounts": solo[solo >= 5].groupby("year").sum(),
})

# arXiv's October 2026 rule: at most two new submissions per account per calendar month.
per_month = S.groupby(["year", "month", "submitter"], observed=True).size()
over_cap = (per_month - 2).clip(lower=0).groupby("year").sum() / per_year["papers"]
accounts_over = (per_month > 2).groupby("year").sum()
per_year.tail(3).round(3)
papers submitters papers per submitter share from 5+ accounts solo share
year
2024 157550 109156 1.443 0.095 0.109
2025 181969 125031 1.455 0.101 0.108
2026 231851 146136 1.587 0.150 0.145

Then the monthly series used in the charts, and the comparison with other venues over February to August.

Show the code
P = papers[(papers.month >= "2015-01") & (papers.month <= LAST)]
monthly = P.groupby("month").agg(
    papers=("id", "size"), solo=("n_authors", lambda s: (s == 1).mean()),
    marker=("llm_marker", "mean"), accepted=("accepted", "mean"), abs_words=("abs_words", "median"))

official_full = official.iloc[:-1]                         # drop the partial current month
def feb_to_m(s):
    s = s[(s.index.month >= 2) & (s.index.month <= M) & (s.index.year >= 2019)]
    return s.groupby(s.index.year).sum()
venues = pd.DataFrame({"arXiv": feb_to_m(official_full), "bioRxiv": feb_to_m(biorxiv),
                       "medRxiv": feb_to_m(medrxiv), "Journal articles (Crossref)": feb_to_m(journals)})
venue_growth = pd.DataFrame({f"{Y0} to {Y1}": venues.loc[Y1] / venues.loc[Y0] - 1,
                             f"{Y1 - 3} to {Y1}": venues.loc[Y1] / venues.loc[Y1 - 3] - 1})

yoy = (n / n.shift(1) - 1)
K = dict(
    n1=f"{n[Y1]:,}", n0=f"{n[Y0]:,}", yoy=f"{yoy[Y1]:.0%}", yoy_prev=f"{yoy[Y0]:.0%}",
    sep=f"{official_full.iloc[-1]:,}", sep_month=official_full.index[-1].strftime("%B %Y"),
    sep_x=f"{official_full.iloc[-1] / official_full.iloc[-25]:.1f}",
    cs_share=f"{g_group.loc['cs', 'share of growth']:.0%}", cs_g=f"{g_group.loc['cs', 'growth']:.0%}",
    math_g=f"{g_group.loc['math', 'growth']:.0%}", phys_g=f"{g_group.loc['physics', 'growth']:.0%}",
    ai_x=f"{by_cat.loc[Y1, 'cs.AI'] / by_cat.loc[Y0, 'cs.AI']:.1f}",
    debut_share=f"{g_tenure.loc['first year on arXiv', 'share of growth']:.0%}",
    vet_share=f"{g_tenure.loc['5+ years', 'share of growth']:.0%}",
    pps1=f"{per_year.loc[Y1, 'papers per submitter']:.2f}", pps0=f"{per_year.loc[Y0, 'papers per submitter']:.2f}",
    solo_g=f"{g_team.loc['1 author', 'growth']:.0%}", solo_share=f"{g_team.loc['1 author', 'share of growth']:.0%}",
    multi_g=f"{(g_team[Y1].sum() - g_team.loc['1 author', Y1]) / (g_team[Y0].sum() - g_team.loc['1 author', Y0]) - 1:.0%}",
    sp1=f"{solo_prolific.loc[Y1, 'accounts with 5+ solo papers']:,}", sp0=f"{solo_prolific.loc[Y0, 'accounts with 5+ solo papers']:,}",
    sp_papers=f"{solo_prolific.loc[Y1, 'papers from 5+ solo accounts']:,}",
    sp_pct=f"{solo_prolific.loc[Y1, 'papers from 5+ solo accounts'] / per_year.loc[Y1, 'papers']:.1%}",
    bio=f"{venue_growth.iloc[1, 0]:.0%}", med=f"{venue_growth.iloc[2, 0]:.0%}", jrn=f"{venue_growth.iloc[3, 0]:.0%}",
    arx=f"{venue_growth.iloc[0, 0]:.0%}",
    cap1=f"{over_cap[Y1]:.1%}", cap0=f"{over_cap[Y0]:.1%}",
    marker_peak=f"{monthly.marker.max():.1%}", marker_peak_m=monthly.marker.idxmax().strftime("%B %Y"),
    marker_now=f"{monthly.marker.iloc[-1]:.1%}",
)
venue_growth.style.format("{:+.1%}")
  2025 to 2026 2023 to 2026
arXiv +27.7% +72.7%
bioRxiv +13.5% +42.0%
medRxiv +21.2% +62.3%
Journal articles (Crossref) +8.0% +27.7%

The short answer

Both stories are partly right, and the metadata shows which part of each.

  1. The lift-off is real, and it is steepest on arXiv. January to August 2026 brought 231,851 new papers, 27% more than the same months a year earlier. Over February to August, bioRxiv grew 13%, medRxiv 21%, and journal articles registered with Crossref 8%.
  2. It is not only an AI-papers story. Computer science supplied 56% of the added papers and cs.AI alone grew 2.5x, but mathematics grew 34% and physics 19% after years of single-digit growth.
  3. It is mostly established accounts submitting more, not a flood of newcomers. Only 13% of the growth came from accounts in their first year on arXiv; 51% came from accounts at least five years old. Papers per account, flat for a decade, jumped from 1.46 to 1.59.
  4. The mill-like signature exists, and it is small. Single-author papers rose 71% (everything else: 22%), reversing a decade-long decline. Accounts posting five or more solo papers in eight months went from 136 to 561. Their solo papers number 4,105, or 1.8% of the total.
  5. Metadata cannot grade quality. It shows output accelerating much faster than peer-reviewed venues are absorbing it. Whether that output is good research is a question for outcomes we can only observe later.

The lift-off is real

Show the code
s = official_full.loc["2008-01":]
x = s.index.to_timestamp()
fig, ax = plt.subplots()
ax.plot(x, s.values, color="#c3c2b7", lw=1, label="Monthly")
ax.plot(x, s.rolling(12).mean().values, color=BLUE, label="12-month average")
ax.scatter(x[-1], s.iloc[-1], s=36, color=BLUE, zorder=3, edgecolor="#fcfcfb", linewidth=2)
ax.annotate(f"{s.index[-1].strftime('%b %Y')}: {s.iloc[-1]:,}", (x[-1], s.iloc[-1]), xytext=(-8, 0),
            textcoords="offset points", ha="right", va="center", color=INK)
ax.yaxis.set_major_formatter(thousands); ax.set_ylim(0); ax.legend(loc="upper left")
ax.set_title("arXiv submissions per month")
plt.show()
Figure 1: New arXiv submissions per month (official arXiv statistics).

Through the 2010s arXiv grew by 5 to 14 percent a year, and it nearly stalled in 2021 and 2022. The first eight months of 2026 grew 27%, after 15% the year before. September 2026, the latest complete month in arXiv’s own statistics, brought 40,363 submissions, 2.0 times the figure of two years earlier. The “lift-off” in the LinkedIn post is not a plotting artefact.

Year-over-year growth for the same January to August window, from the per-paper table:

Show the code
pd.DataFrame({"papers": n, "growth": yoy}).loc[2016:].style.format({"papers": "{:,}", "growth": "{:+.1%}"})
  papers growth
year    
2016 73,427 +8.5%
2017 79,435 +8.2%
2018 90,756 +14.3%
2019 100,629 +10.9%
2020 116,077 +15.4%
2021 119,096 +2.6%
2022 120,505 +1.2%
2023 133,745 +11.0%
2024 157,550 +17.8%
2025 181,969 +15.5%
2026 231,851 +27.4%

Where the growth comes from

Show the code
fig, axes = plt.subplots(1, 2, figsize=(9.2, 4.2), gridspec_kw={"wspace": 0.55})
names = {"cs": "Computer science", "physics": "Physics", "math": "Mathematics", "stat": "Statistics",
         "econ": "Economics", "eess": "Electrical eng.", "q-fin": "Quant. finance", "q-bio": "Quant. biology"}
for ax, t, title, lab in [(axes[0], g_group, "By subject group", lambda i: names[i]),
                          (axes[1], g_cat.head(10), "Top 10 categories", str)]:
    t = t.iloc[::-1]
    ax.barh([lab(i) for i in t.index], t["added"], color=BLUE, height=0.62)
    for y, (a, g) in enumerate(zip(t["added"], t["growth"])):
        ax.text(max(a, 0), y, f"  {a:+,} ({g:+.0%})", va="center", ha="left", color=INK2, fontsize=9)
    ax.set_title(title); ax.grid(False); ax.set_xticks([]); ax.spines["bottom"].set_visible(False)
    ax.set_xlim(min(0, t["added"].min()), t["added"].max() * 1.75); ax.tick_params(length=0)
fig.suptitle(f"Papers added, Jan to {LAST.strftime('%b')} {Y0} vs {Y1} (growth in brackets)", x=0.02, ha="left", fontsize=10, color=INK2)
plt.show()
Figure 2: Papers added between the two January to August windows, by subject group and by primary category.

Computer science accounts for 56% of the added papers, and one category stands out: cs.AI grew 2.5x in a single year. arXiv singled out the same category when it introduced submission rate limits on 1 October 2026.

The more surprising result is outside CS. Mathematics grew 34% and physics 19%, in fields that had been growing by a few percent a year. Combinatorics (math.CO) grew faster than computer vision. A surge that reaches pure mathematics is not explained by “more people are working on AI”. Something changed in how quickly papers get written across fields.

The full group table is in Step 6. The twenty categories that added the most papers:

Show the code
g_cat.head(20).style.format({"growth": "{:+.1%}", "share of growth": "{:.1%}", Y0: "{:,}", Y1: "{:,}", "added": "{:+,}"})
  2025 2026 added growth share of growth
primary          
cs.AI 4,920 12,136 +7,216 +146.7% 14.5%
cs.LG 16,484 21,456 +4,972 +30.2% 10.0%
cs.CV 17,837 21,769 +3,932 +22.0% 7.9%
quant-ph 7,415 10,017 +2,602 +35.1% 5.2%
cs.RO 5,051 7,500 +2,449 +48.5% 4.9%
math.CO 2,775 4,417 +1,642 +59.2% 3.3%
cs.CR 3,684 5,217 +1,533 +41.6% 3.1%
cs.SE 2,593 3,856 +1,263 +48.7% 2.5%
cs.CL 11,593 12,811 +1,218 +10.5% 2.4%
math.AP 3,112 4,136 +1,024 +32.9% 2.1%
stat.ME 2,052 2,964 +912 +44.4% 1.8%
math.OC 2,738 3,427 +689 +25.2% 1.4%
math.NA 2,361 2,976 +615 +26.0% 1.2%
eess.SY 2,639 3,252 +613 +23.2% 1.2%
hep-th 2,494 3,100 +606 +24.3% 1.2%
astro-ph.IM 1,004 1,603 +599 +59.7% 1.2%
math.NT 1,792 2,386 +594 +33.1% 1.2%
cond-mat.mtrl-sci 3,769 4,356 +587 +15.6% 1.2%
cs.IT 1,491 2,037 +546 +36.6% 1.1%
math.AG 1,497 2,040 +543 +36.3% 1.1%

Who is submitting more

Show the code
fig, axes = plt.subplots(1, 2, figsize=(9.2, 3.8), gridspec_kw={"wspace": 0.3})
ax = axes[0]
pps = per_year["papers per submitter"]
ax.plot(pps.index, pps.values, color=BLUE, marker="o", ms=5, markeredgecolor="#fcfcfb")
ax.annotate(f"{pps.iloc[-1]:.2f}", (pps.index[-1], pps.iloc[-1]), xytext=(-6, 4), textcoords="offset points", ha="right")
ax.set_ylim(1.3, 1.65); ax.set_title("Papers per account"); ax.set_xticks(range(2015, Y1 + 1, 2))
ax = axes[1]
ax.bar(g_tenure.index.astype(str), g_tenure["share of growth"], color=BLUE, width=0.55)
for i, v in enumerate(g_tenure["share of growth"]):
    ax.text(i, v + 0.01, f"{v:.0%}", ha="center", color=INK)
ax.yaxis.set_major_formatter(pct0); ax.set_ylim(0, 0.6); ax.set_title(f"Share of {Y0} to {Y1} growth, by account age")
plt.show()
Figure 3: Left: papers per submitting account within each January to August window. Right: which accounts supplied the growth.

If paper mills or first-time hobbyists were driving the surge, new accounts would dominate the growth. They do not. Accounts in their first year supplied 13% of the added papers, and their output grew more slowly than anyone else’s, which fits arXiv’s tightened endorsement rules for new submitters in January 2026. Half of the growth (51%) came from accounts that have been on arXiv for five years or more.

What changed is the pace. For ten years the average account submitted about 1.45 papers in an eight-month window, almost without variation. In 2026 that rose to 1.59. The people who were already writing papers are writing more of them.

The three splits behind this section. Accounts that submitted five or more papers in the window are where output grew fastest:

Show the code
fmt = {"growth": "{:+.1%}", "share of growth": "{:.1%}", Y0: "{:,}", Y1: "{:,}", "added": "{:+,}"}
pd.concat({"Account age": g_tenure, "Authors on the paper": g_team, "Papers by the account in the window": g_volume}).style.format(fmt)
    2025 2026 added growth share of growth
Account age first year on arXiv 47,734 54,218 +6,484 +13.6% 13.0%
1-4 years 57,908 76,023 +18,115 +31.3% 36.3%
5+ years 76,327 101,610 +25,283 +33.1% 50.7%
Authors on the paper 1 author 19,575 33,511 +13,936 +71.2% 27.9%
2-3 66,143 83,856 +17,713 +26.8% 35.5%
4-6 61,751 72,016 +10,265 +16.6% 20.6%
7+ 34,500 42,468 +7,968 +23.1% 16.0%
Papers by the account in the window 1 paper 92,469 102,144 +9,675 +10.5% 19.4%
2 42,672 53,604 +10,932 +25.6% 21.9%
3-4 28,469 41,386 +12,917 +45.4% 25.9%
5-9 12,970 23,446 +10,476 +80.8% 21.0%
10+ 5,389 11,271 +5,882 +109.1% 11.8%

The solo-author break

Show the code
fig, axes = plt.subplots(1, 2, figsize=(9.2, 3.8), gridspec_kw={"wspace": 0.3})
ax = axes[0]
sm = monthly["solo"].rolling(3).mean()
ax.plot(sm.index.to_timestamp(), sm.values, color=BLUE)
ax.yaxis.set_major_formatter(pct0); ax.set_ylim(0, 0.25); ax.set_title("Single-author share (3-month average)")
ax = axes[1]
sp = solo_prolific["accounts with 5+ solo papers"]
ax.bar(sp.index, sp.values, color=BLUE, width=0.62)
ax.text(sp.index[-1], sp.iloc[-1] + 8, f"{sp.iloc[-1]:,}", ha="center"); ax.text(sp.index[-2], sp.iloc[-2] + 8, f"{sp.iloc[-2]:,}", ha="center")
ax.set_title("Accounts with 5+ solo papers"); ax.set_xticks(range(2015, Y1 + 1, 2))
plt.show()
Figure 4: Left: share of new papers with a single author, by month. Right: accounts that submitted five or more single-author papers in a January to August window.

This is the clearest structural break in the data. Research had been getting steadily more collaborative: the single-author share of new papers fell from about 20 percent in 2015 to under 11 percent in 2025. In 2026 it turned around. Single-author papers grew 71% while multi-author papers grew 22%, and solo papers alone account for 28% of all the growth.

At the extreme end is the pattern Srijit described. The number of accounts posting five or more single-author papers in eight months sat around 100 for a decade, reached 136 in 2025, and then 561 in 2026. One person producing a paper every few weeks, alone, was rare enough to be a curiosity. It no longer is.

It is also a small part of the whole. Their solo papers number 4,105, or 1.8% of the window. arXiv’s new cap of two submissions per account per month would have held back 4.0% of 2026’s papers (it was 2.6% the year before). The high-frequency tail is real and growing fast, but removing it entirely would leave nearly all of the surge in place.

Show the code
per_year.join(solo_prolific).join(over_cap.rename("share above 2/month")).style.format({
    "papers": "{:,}", "submitters": "{:,}", "papers per submitter": "{:.3f}", "share from 5+ accounts": "{:.1%}",
    "solo share": "{:.1%}", "papers from 5+ solo accounts": "{:,}", "share above 2/month": "{:.1%}"})
  papers submitters papers per submitter share from 5+ accounts solo share accounts with 5+ solo papers accounts with 10+ solo papers papers from 5+ solo accounts share above 2/month
year                  
2015 67,695 46,222 1.465 8.4% 20.5% 86 12 722 3.0%
2016 73,427 50,506 1.454 8.5% 19.3% 86 14 792 3.0%
2017 79,435 54,910 1.447 8.3% 18.1% 87 12 747 3.0%
2018 90,754 62,499 1.452 8.6% 16.3% 92 17 851 3.0%
2019 100,628 70,078 1.436 8.2% 14.9% 76 11 658 2.7%
2020 116,077 79,849 1.454 8.9% 14.1% 121 8 845 2.6%
2021 119,096 83,489 1.426 8.4% 13.2% 101 12 751 2.5%
2022 120,505 85,326 1.412 8.1% 12.7% 90 9 747 2.5%
2023 133,744 94,569 1.414 8.3% 12.2% 89 13 745 2.3%
2024 157,550 109,156 1.443 9.5% 10.9% 88 13 762 2.8%
2025 181,969 125,031 1.455 10.1% 10.8% 136 17 1,035 2.6%
2026 231,851 146,136 1.587 15.0% 14.5% 561 75 4,105 4.0%

No accounts are named here on purpose. A submitter name is not proof of anything: prolific legitimate accounts exist (proceedings series, large collaborations), and common names merge several people.

Are credible venues growing the same way?

Show the code
base = Y1 - 3
idx = venues.loc[base:] / venues.loc[base] * 100
fig, ax = plt.subplots(figsize=(7.6, 4.0))
for col, c in zip(idx.columns, [BLUE, ORANGE, AQUA, YELLOW]):
    ax.plot(idx.index, idx[col], color=c, marker="o", ms=5, markeredgecolor="#fcfcfb", label=col)
    ax.annotate(f"{col}  {idx[col].iloc[-1]:.0f}", (Y1, idx[col].iloc[-1]), xytext=(8, 0), textcoords="offset points", va="center", color=INK)
ax.set_xticks(idx.index); ax.set_xlim(base - 0.15, Y1 + 1.6); ax.set_ylim(95, None)
ax.legend(loc="upper left"); ax.set_title(f"Growth since {base} ({base} = 100)")
plt.show()
Figure 5: Output over February to August of each year, indexed to 2023. Crossref counts are journal articles by publication date.

This was the test proposed in the comments, and the answer is no. Over three years arXiv grew about 73%, while journal articles registered with Crossref grew 28%. In the latest year arXiv grew 28% against 8% for journals. The life-science preprint servers sit in between. medRxiv comes closest, at 21% in the latest year, but it is a twentieth of arXiv’s size and was already growing almost that fast; neither server shows a jump in 2026 like arXiv’s.

Two cautions. Crossref counts every journal with a DOI, so “credible, defined broadly” is doing a lot of work, and journal growth includes paper-mill output of its own. And journals lag preprints: a paper posted in 2026 may be published in 2027. A gap this wide still means that preprints are being produced far faster than the peer-review system is currently absorbing them. Either a wave of publications is on its way, or a growing share of arXiv will never be reviewed by anyone.

Show the code
venues.style.format("{:,}")
  arXiv bioRxiv medRxiv Journal articles (Crossref)
month        
2019 88,954 16,587 261 2,160,504
2020 104,080 23,692 9,093 2,345,416
2021 106,332 21,940 7,814 2,590,402
2022 106,891 20,398 6,276 2,669,359
2023 119,871 22,242 6,414 2,782,909
2024 139,994 24,834 7,369 3,006,477
2025 162,188 27,831 8,589 3,290,339
2026 207,036 31,575 10,408 3,552,626

What the abstracts can and cannot tell us

Show the code
mk = monthly["marker"].loc["2019-01":]
fig, ax = plt.subplots()
ax.plot(mk.index.to_timestamp(), mk.values, color=BLUE)
ax.axvline(pd.Timestamp("2022-11-30"), color=MUTED, lw=1, ls=(0, (3, 3)))
ax.text(pd.Timestamp("2022-11-01"), mk.max(), "ChatGPT released ", ha="right", va="top", color=INK2)
ax.yaxis.set_major_formatter(pct0); ax.set_ylim(0); ax.set_title("Abstracts with LLM marker words")
plt.show()
Figure 6: Share of new abstracts using at least one of: delve, underscore, showcase, pivotal, intricate, meticulous (any inflection).

An obvious idea is to detect AI-written text directly. Researchers have done this through “excess vocabulary”: words such as delve that became abruptly more common after ChatGPT appeared (Kobak et al.). The same handful of words in arXiv abstracts shows the rise clearly, peaking at 7.2% of abstracts in April 2024.

Then it falls away, back to 1.2% by August 2026, during the exact period when submissions took off. Nobody believes LLM use declined in 2026. Models changed their habits and authors learned to edit out the tells. A fixed word list is a detector for one generation of models, and it says nothing about quality in any case: a careful paper polished by an LLM and a fabricated one both trip it.

Abstracts did get longer, though: the median went from 166 words in early 2025 to 179 by mid 2026, after creeping up by a word or two a year.

Why publication status cannot settle it yet

The natural quality measure is whether a preprint ends up published. The snapshot has three signals: a journal reference, a DOI, and comments mentioning acceptance. All three are added after the fact, so recent cohorts always look worse than old ones regardless of quality:

Show the code
c = W.groupby("year").agg(journal_ref=("has_jref", "mean"), doi=("has_doi", "mean"), accepted_in_comments=("accepted", "mean"))
c.loc[2018:].style.format("{:.1%}")
  journal_ref doi accepted_in_comments
year      
2018 32.6% 48.4% 21.4%
2019 31.2% 45.9% 20.7%
2020 28.7% 43.8% 21.2%
2021 25.3% 42.2% 21.3%
2022 21.7% 39.5% 21.5%
2023 20.6% 31.0% 21.2%
2024 18.4% 24.4% 20.4%
2025 15.2% 20.0% 18.6%
2026 6.4% 8.5% 14.1%

The decline starts years before the surge and steepens toward the present, which is what right-censoring looks like. Comparing cohorts fairly needs each one measured at the same age, which means keeping successive snapshots. That is the follow-up worth doing: re-run this notebook in a year and compare the 2026 cohort’s publication rate at twelve months with the 2025 cohort’s today.

So: acceleration or slop?

The data rule out the simple version of each story.

It is not a flood of outsiders and mills. The growth comes mainly from people with a track record on arXiv, across computer science, mathematics and physics at once, and the accounts that behave like mills are a few percent of the total at most.

It is also not business as usual at higher speed. Papers per account jumped after a decade of stability, solo work reversed a long decline, and peer-reviewed output did not move with it. The same researchers are producing papers faster than before, increasingly alone, and faster than journals and conferences are certifying them.

Whether that is acceleration of research or only of papers is not something titles, author lists and abstracts can decide. arXiv’s moderators, who read the submissions, describe a rise in “thin papers of narrow scope” and work split into smaller pieces, and have responded with endorsement rules and rate limits. The measurements that would settle it are outcomes: what share of the 2026 cohort is published a year on, how it is cited, and how often it is withdrawn. Those can be measured, just not yet.

Limitations

  • Submitter names are not identities. The snapshot has display names, not account IDs. Common names merge several people and inflate “papers per account”; a person who changes their display name is split. This affects every year, but collisions grow with volume, so the level of the prolific-account counts is less reliable than their direction. The solo-author share does not depend on names at all.
  • Account age is measured from the first paper under that name in the snapshot, with the same caveat.
  • The comparison window is January to August. September 2026 appears only in the official monthly count; the per-paper snapshot was taken before that month ended.
  • Crossref dates are publication dates as deposited by publishers. Year-only dates land on 1 January (excluded here), and recent months can still gain records as publishers deposit late, which would slightly understate journal growth.
  • Primary category only. Cross-listed papers are counted once, under their first category.
  • No full text and no moderation data. Rejected submissions never appear in the snapshot, so this is the growth that survived arXiv’s moderation.

Reproducing this

conda env create -f environment.yml && conda activate arxiv-analysis
quarto render arxiv-liftoff.qmd

The first render downloads about 1.8 GB and takes several minutes to parse; later renders reuse data/ and finish in seconds. Delete data/raw/ to refresh every source.