Contents
  1. A model label can be forged
  2. Evading a source is not reaching a target
  3. How ForgePrint works
  4. Results
  5. Which target you aim at matters more than where you start
  6. From an open model to a chosen commercial one
  7. Case study
  8. Takeaways
  9. Ethics and release
  10. BibTeX
  11. References

Forging LLM Authorship Fingerprints with Targeted Rewriting

Haohan Yuan
UNC Charlotte
Simin Chen
George Mason University
Xi Niu
UNC Charlotte
Hanqing Guo
Indiana University
Depeng Xu
UNC Charlotte
Haopeng Zhang *
UNC Charlotte
*Corresponding author

A model label can be forged

Classifiers can often identify which LLM produced an unmodified response: ours reach 85.9% accuracy. But a high-confidence model label is not proof of who or what originally generated the text.

ForgePrint rewrites an LLM response to preserve its meaning while making it look like it came from a chosen target model. Across four domains, it achieves a 60.8% attack success rate. On CNN/DailyMail, it reaches 70.2%.

In short: text can be made to look like it was written by one model even when it was originally generated by another.

Starting from a Claude summary, existing rewriters largely preserve the Claude fingerprint, whereas ForgePrint shifts attribution toward the chosen GPT fingerprint.
Figure 1. The task: targeted rewriting. Put through existing rewriters, a Claude summary still reads as Claude. ForgePrint moves it to GPT — the model we chose.

Evading a source is not reaching a target

Most work on fooling these detectors only asks the text to stop looking like its real source — any other answer counts as a win. Landing on the one model you picked is much harder.

The two goals sound similar and score very differently. Planner TST and ForgePrint-4B are 10.5 points apart at simply escaping the source, but 27.7 points apart at reaching the intended target (33.1% vs. 60.8%).

A GPT summary rewritten toward Gemini is classified as Grok: source evasion succeeds, but targeted transfer fails.
Figure 2. Escaping is not arriving. A GPT summary rewritten toward Gemini comes out looking like Grok: it left its real source, but missed the model it was aimed at.

How ForgePrint works

No corpus of paired rewrites exists — one model’s text rewritten to look like another’s — so ForgePrint builds its own. For each source→target pair it assembles editing operators (drop a habit of the source, adopt one of the target’s, reshape the sentences) and keeps only those that measurably move a surrogate classifier, a stand-in it trains itself, toward the target. A large Teacher applies them, generates several candidate rewrites and selects the best. A 4B Student is distilled from those, then refined on its own output, ranked by factual accuracy first and target resemblance second, so it cannot succeed by quietly altering what the text says.

Only the Student is deployed. Given the text and the two model names it rewrites in a single pass — no operator bank, no surrogate, no access to the original article.

ForgePrint pipeline: an offline operator bank, a Teacher that generates and selects candidate rewrites, and a Student distilled with SFT, DPO, and GRPO.
Figure 3. How the rewriter is built.

Results

The attack never sees what grades it. It is developed against its own surrogate, trained on one half of the data. Every number below comes from four held-out classifiers — RoBERTa, DeBERTa, GPT-2 and TF-IDF, trained on the other half, which shares no documents with the first — and the attack never queries them.

Table 1: main results on targeted fingerprint transfer across four domains, comparing six published baselines against the ForgePrint Teacher and the 4B Student.
Table 1. Main results, on all four kinds of text. SrcEv is how often the rewrite escapes its real source, ASR how often it reaches the target we picked, Align how well it still matches the source document. Macro₄ averages the four; the small ± figures are 95% confidence half-widths. Method names link to the references.

The 4B Student beats the 26B Teacher it learned from (70.2% vs. 54.1% on news), because its last training stage lets it practise on its own rewrites instead of only copying the Teacher. The ordering holds across all four kinds of text — news, research papers, dialogue and how-to guides — and all four classifiers.

The gain is not a loss of faithfulness

A rewrite that quietly altered the content would score well and mean nothing. Compared only against rewrites of matched faithfulness, the lead over the strongest published method holds at 27.5 points under AlignScore and 28.3 under MiniCheck.

▸Table 2 — the faithfulness-adjusted numbers
Table 2: raw and faithfulness-adjusted Macro-4 target ASR across four domains, using AlignScore or MiniCheck to match rewrite faithfulness.
Table 2. Success rate before and after holding faithfulness constant.
PCA projection through a held-out GPT-2 evaluator showing rewrites moving from the source region toward the chosen target region.
Figure 4. A map of what the classifier sees. Each summary is placed by the writing features a held-out classifier reads from it, flattened to two dimensions (PCA); a line joins each summary to its rewrite. The rewrites do not merely drift away from the source — they travel toward the target. The share landing on the intended model rises from 44% to 66% to 78%.

Which target you aim at matters more than where you start

Claude is the hardest model to imitate, reached 39.7% of the time on average against 64.1–70.3% for the others — and how hard it is depends on what is being written: aiming at Claude works 5% of the time on research-paper summaries and 70% on news. Which model you aim at accounts for 82.6% of the difference between one rewriting job and another; which model you start from accounts for the other 17.4%.

Heatmap of target ASR by target model and domain: Claude is hardest at 39.7% mean, Grok easiest at 70.3%.
Figure 5. How often the rewrite reaches each target, by kind of text.

From an open model to a chosen commercial one

Here an open model, Gemma, writes the news summaries, and the rewrite aims them at one named commercial model. Left alone, a Gemma summary is almost never mistaken for Claude (0.1%) or Grok (0.5%). After rewriting, it passes for the chosen model between 49.6% and 81.5% of the time.

Targeted success per path (%) ForgePrint Teacher ForgePrint-4B
Gemma → Claude
35.8
49.6
Gemma → GPT
17.3
64.6
Gemma → Gemini
50.1
77.3
Gemma → Grok
36.1
81.5
Macro
34.8
68.3
0255075100

Here the classifier can also answer “Gemma”, so a rewrite has to leave its real source and reach the model it was aimed at. Not on the same scale as Table 1.

The Student beats its Teacher on three of the four targets, most sharply on GPT (64.6% vs. 17.3%). A Student built on a different backbone, Qwen3.5-9B, reaches 66.7% on average, so this is not specific to the Gemma family.

Case study

CNN/DMGrok → Gemini
cnndm_CorpusB:22:Grok_to_Gemini

Manchester United is preparing for its upcoming UEFA U19 Youth League campaign by conducting trials of promising young players. Specifically, the club will assess 18-year-old defender Luke Tingey and 18-year-old midfielder Kyran Wiltshire. Both players are members of the MK Dons U18 squad that recently won the Youth Alliance South Cup. This trial period will involve training sessions at Carrington.

Held-out evaluators — 3 of 4 say Gemini
RoBERTaGemini
DeBERTaGemini
GPT-2Gemini
TF-IDFGrok

Turns two long sentences into four short ones with one point each, opens with a framing sentence, and drops the free-kick distance and the adjective “lively”. Three of the four evaluators now name Gemini.

The Student is not copying Gemini’s words but the way it arranges them: two long sentences become four short ones making one point each, a framing sentence opens the piece, and incidental detail is dropped. The strongest published method instead reaches for exclamations and evaluative adjectives — and leaves behind a stray bracket from its own prompt template.

Takeaways

Ethics and release

Targeted rewriting is dual-use — the capability that measures the weakness is the one that exploits it — so we keep the experiments contained: everything runs offline on generated summaries and offline classifiers, we query commercial APIs only to produce summaries, and we never test impersonation against a deployed service.

Our code and evaluation suite will be released at github.com/HaohanYuan01/ForgePrint. The trained rewriter is available on request under the terms of the Ethics Statement.

BibTeX

@misc{yuan2026forgeprint,
title = {Forging LLM Authorship Fingerprints with Targeted Rewriting},
author = {Yuan, Haohan and Chen, Simin and Niu, Xi and Guo, Hanqing and Xu, Depeng and Zhang, Haopeng},
year = {2026},
eprint = {2609.38831},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}

References

David, I., & Gervais, A. (2025). Authormist: Evading ai text detectors with reinforcement learning. arXiv Preprint arXiv:2503.08716.
Gong, H., Bhat, S., Wu, L., Xiong, J., & Hwu, W. (2019). Reinforcement Learning Based Text Style Transfer without Parallel Training Corpus. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT).
Horvitz, Z., Patel, A., Singh, K., Callison-Burch, C., McKeown, K., & Yu, Z. (2024). TinyStyler: Efficient few-shot text style transfer with authorship embeddings. Findings of the Association for Computational Linguistics: EMNLP 2024, 13376–13390.
Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (2023). Paraphrasing Evades Detectors of AI-Generated Text, but Retrieval is an Effective Defense. Advances in Neural Information Processing Systems, 36, 27469–27500.
Reif, E., Ippolito, D., Yuan, A., Coenen, A., Callison-Burch, C., & Wei, J. (2022). A Recipe for Arbitrary Text Style Transfer with Large Language Models. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 837–848.
Suzgun, M., Melas-Kyriazi, L., & Jurafsky, D. (2022). Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2195–2222.
Zhang, L., Chuang, Y.-N., Wang, G., Tang, R., Cai, X., Shenoy, R., & Hu, X. (2025). A Decoupled Multi-Agent Framework for Complex Text Style Transfer. Findings of the Association for Computational Linguistics: EMNLP 2025, 21393–21403.