# Large Language Models# LM Arena# LLM identification
Abstract
Voting-based leaderboards such as LM Arena are central to open-ended LLM evaluation, but their reliability depends on response anonymity. Prior work shows that surface features such as bag-of-words and TF-IDF can reveal model identity, yet such signals are unreliable when models are stylistically similar or closely matched, and are easily down-weighted during vote aggregation. To examine the remaining anonymity risk, we introduce InterPol, a model-driven identification framework that learns target-specific response signatures from interpolated preference data. InterPol synthesizes controllable hard negatives by interpolating a target-mimicking copy model with its initial backbone, and trains a target-specific detector via curriculum-based triplet preference learning. Across three frontier target LLMs and two query sources, InterPol yields average relative gains of 32% in Accuracy and 25% in AUROC over prior methods. Moreover, in arena-style promotion simulations, InterPol attains the same ranking gain with about 19% fewer interactions than existing identification baselines; with selective competitor downvoting, interactions are further reduced by about 68% relative to target-only voting.
Motivation
LM Arena evaluates LLMs through anonymous pairwise comparisons: users see two responses to the same query and vote for the better answer without knowing which models generated them.
If response anonymity breaks, a strategic voter can infer which model produced each answer and cast only the votes that promote a target or penalize its competitors—directly threatening the reliability of the resulting ranking.
LM ArenaAnonymous battle
Model A?
Model B?
Which response is better?
A is betterTieB is better
Existing approaches reveal the risk—but rely on mitigable surface cues
Identification from surface cues [1]
Simple features such as response length, bag-of-words, and TF-IDF can identify LLM responses with high accuracy. However, defenders can exploit the same signals to down-weight suspicious votes, and surface cues become less reliable for closely related or stylistically similar models.
Identification enables strategic voting [2]
Identifying a target model enables selective vote rigging, but target-only voting is interaction-inefficient because the target appears in only about 1% of sampled battles. An omnipresent strategy improves efficiency by exploiting dependencies in the Elo ranking system.
“Do more advanced identification techniques exist that could threaten the robustness and credibility of leaderboards?”
Threat model. We consider a model developer seeking to raise its own model's rank while participating as a normal leaderboard user. The attacker builds a target-specific detector from responses collected through public model access and, during a battle, uses only the anonymous response text—without metadata, model labels, or access to system internals.
Method
Overview of InterPol. A copy model learns to mimic the target LLM. Interpolating it with the initial backbone produces target-like negatives of controllable difficulty for adaptive detector training.
1
Select a real hard negative
For each query, collect target and competing responses, then choose the non-target response most semantically similar to the target.
2
Synthesize interpolated negatives
Fine-tune a copy model to imitate the target, then interpolate it with its initial backbone to control negative difficulty.
3
Learn with an adaptive curriculum
Train the detector to score target responses above interpolated responses and interpolated responses above other models.
1. Real hard-negative selection
Among responses from non-target models, InterPol selects the response closest to the target in sentence-embedding space:
The interpolated model generates $\widetilde{y}_{\alpha}^{i}=\widetilde{M}_{\alpha}(q^i)$. As $\alpha$ increases, its responses move toward the target-mimicking copy model and become harder negatives.
3. Triplet preference learning
Each training item is $(q^i,y_*^i,\widetilde{y}_{\alpha}^{i},y_o^i)$, ordered as target $\succ$ interpolated $\succ$ non-target. We write the Bradley–Terry term in its negative-log form so that every expression below is a loss to minimize:
If $m^i<0$, the detector falls back to the easier pairwise objective for that sample; otherwise it uses the full triplet objective. Training proceeds through two stages, $\alpha_1=0.5$ followed by the harder $\alpha_2=0.75$.
Experiments
[ Main identification results ]
InterPol achieves the best overall performance, outperforming all identification baselines across every target LLM and query source! 🏆
[ Score distributions of the detector ]
The initial detector cannot reliably distinguish target from non-target responses. Triplet training at $\alpha=0.5$ begins to separate them, and iterative curriculum learning at $\alpha=0.75$ widens the score margin further—producing a clear target-specific decision boundary.
Across targets starting at Ranks 5, 15, and 45, InterPol reaches the same promotion goals with the lowest or competitive number of Arena interactions.
[ Competitor-aware voting creates more useful battles ]
For a target starting at Rank 5, competitor-aware voting reduces the required interactions by 69.9% on average across promotion goals from Rank 4 to Rank 1.
Citation
BibTeX
@inproceedings{cho2026interpol,
title = {InterPol: De-anonymizing LM Arena via Interpolated Preference Learning},
author = {Cho, Minsung and Kim, Jaehyung},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}