EMNLP 2026 MAIN

InterPol: De-anonymizing LM Arena via Interpolated Preference Learning

Minsung Cho1  ·  Jaehyung Kim1
1Yonsei University
# Large Language Models # LM Arena # LLM identification

Abstract

Voting-based leaderboards such as LM Arena are central to open-ended LLM evaluation, but their reliability depends on response anonymity. Prior work shows that surface features such as bag-of-words and TF-IDF can reveal model identity, yet such signals are unreliable when models are stylistically similar or closely matched, and are easily down-weighted during vote aggregation. To examine the remaining anonymity risk, we introduce InterPol, a model-driven identification framework that learns target-specific response signatures from interpolated preference data. InterPol synthesizes controllable hard negatives by interpolating a target-mimicking copy model with its initial backbone, and trains a target-specific detector via curriculum-based triplet preference learning. Across three frontier target LLMs and two query sources, InterPol yields average relative gains of 32% in Accuracy and 25% in AUROC over prior methods. Moreover, in arena-style promotion simulations, InterPol attains the same ranking gain with about 19% fewer interactions than existing identification baselines; with selective competitor downvoting, interactions are further reduced by about 68% relative to target-only voting.

Motivation

LM Arena evaluates LLMs through anonymous pairwise comparisons: users see two responses to the same query and vote for the better answer without knowing which models generated them.

If response anonymity breaks, a strategic voter can infer which model produced each answer and cast only the votes that promote a target or penalize its competitors—directly threatening the reliability of the resulting ranking.

Existing approaches reveal the risk—but rely on mitigable surface cues

Identification from surface cues [1]

Simple features such as response length, bag-of-words, and TF-IDF can identify LLM responses with high accuracy. However, defenders can exploit the same signals to down-weight suspicious votes, and surface cues become less reliable for closely related or stylistically similar models.

Identification enables strategic voting [2]

Identifying a target model enables selective vote rigging, but target-only voting is interaction-inefficient because the target appears in only about 1% of sampled battles. An omnipresent strategy improves efficiency by exploiting dependencies in the Elo ranking system.

The Remaining Question
“Do more advanced identification techniques exist that could threaten the robustness and credibility of leaderboards?”
Threat model. We consider a model developer seeking to raise its own model's rank while participating as a normal leaderboard user. The attacker builds a target-specific detector from responses collected through public model access and, during a battle, uses only the anonymous response text—without metadata, model labels, or access to system internals.

Method

InterPol pipeline: negative synthesis through model interpolation and adaptive curriculum learning
Overview of InterPol. A copy model learns to mimic the target LLM. Interpolating it with the initial backbone produces target-like negatives of controllable difficulty for adaptive detector training.
1

Select a real hard negative

For each query, collect target and competing responses, then choose the non-target response most semantically similar to the target.

2

Synthesize interpolated negatives

Fine-tune a copy model to imitate the target, then interpolate it with its initial backbone to control negative difficulty.

3

Learn with an adaptive curriculum

Train the detector to score target responses above interpolated responses and interpolated responses above other models.

1. Real hard-negative selection

Among responses from non-target models, InterPol selects the response closest to the target in sentence-embedding space:

\[ y_o^i=\underset{y_k^i\in\mathcal{Y}^i}{\arg\max}\; \operatorname{sim}\!\left(e(y_*^i),e(y_k^i)\right). \]

2. Copy-model training and interpolation

A copy model is first trained on target responses by supervised fine-tuning:

\[ \widehat{M}_*= \underset{\phi}{\arg\min}\;\frac{1}{|\mathcal{D}|} \sum_{(q^i,y_*^i)\in\mathcal{D}} \mathcal{L}_{\mathrm{CE}}\!\left(\widehat{M}_{\phi}(q^i),y_*^i\right). \]

Its trained weights are then interpolated with the initial backbone:

\[ \widetilde{\phi}_{\alpha}=(1-\alpha)\phi_{\mathrm{init}}+\alpha\widehat{\phi}_{\mathrm{copy}}, \qquad \alpha\in[0,1] \]

The interpolated model generates $\widetilde{y}_{\alpha}^{i}=\widetilde{M}_{\alpha}(q^i)$. As $\alpha$ increases, its responses move toward the target-mimicking copy model and become harder negatives.

3. Triplet preference learning

Each training item is $(q^i,y_*^i,\widetilde{y}_{\alpha}^{i},y_o^i)$, ordered as target $\succ$ interpolated $\succ$ non-target. We write the Bradley–Terry term in its negative-log form so that every expression below is a loss to minimize:

\[ \ell_{\mathrm{pref}}(q,y_w,y_l)= -\log\sigma\!\left(f_\theta(q,y_w)-f_\theta(q,y_l)\right). \]

The final objective averages the three pairwise relations in each triplet:

\[ \begin{aligned} \mathcal{L}_{\mathrm{fin}} &=\frac{1}{|\mathcal{T}_{\alpha}|} \sum_{i\in\mathcal{T}_{\alpha}}\mathcal{L}_{\mathrm{trp}}^i,\\ \mathcal{L}_{\mathrm{trp}}^i &=\lambda_1\ell_{\mathrm{pref}}(q^i,y_*^i,\widetilde{y}_{\alpha}^{i})\\ &\quad+\lambda_2\ell_{\mathrm{pref}}(q^i,\widetilde{y}_{\alpha}^{i},y_o^i) +\lambda_3\ell_{\mathrm{pref}}(q^i,y_*^i,y_o^i). \end{aligned} \]

Adaptive curriculum

InterPol checks whether the detector already ranks the target above the real hard negative:

\[ m^i=f_\theta(q^i,y_*^i)-f_\theta(q^i,y_o^i). \]

If $m^i<0$, the detector falls back to the easier pairwise objective for that sample; otherwise it uses the full triplet objective. Training proceeds through two stages, $\alpha_1=0.5$ followed by the harder $\alpha_2=0.75$.

Experiments

[ Main identification results ]
Main target-present identification results across GPT-4o, Gemini-Pro, and Claude-4 on Arena and Alpaca
InterPol achieves the best overall performance, outperforming all identification baselines across every target LLM and query source! 🏆
[ Score distributions of the detector ]
Detector score distributions before training, after triplet training at alpha 0.5, and after iterative training at alpha 0.75
The initial detector cannot reliably distinguish target from non-target responses. Triplet training at $\alpha=0.5$ begins to separate them, and iterative curriculum learning at $\alpha=0.75$ widens the score margin further—producing a clear target-specific decision boundary.

Simulations

[ Identification accuracy reduces interaction cost ]
Table 4: interactions and votes required to promote three target models under different identification methods
Across targets starting at Ranks 5, 15, and 45, InterPol reaches the same promotion goals with the lowest or competitive number of Arena interactions.
[ Competitor-aware voting creates more useful battles ]
Table 5: effect of competitor-aware voting on interactions required to promote the Rank 5 target model
For a target starting at Rank 5, competitor-aware voting reduces the required interactions by 69.9% on average across promotion goals from Rank 4 to Rank 1.
Citation

BibTeX

@inproceedings{cho2026interpol,
  title     = {InterPol: De-anonymizing LM Arena via Interpolated Preference Learning},
  author    = {Cho, Minsung and Kim, Jaehyung},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}