Forging LLM Authorship Fingerprints with Targeted Rewriting
A model label can be forged
Classifiers can often identify which LLM produced an unmodified response: ours reach 85.9% accuracy. But a high-confidence model label is not proof of who or what originally generated the text.
ForgePrint rewrites an LLM response to preserve its meaning while making it look like it came from a chosen target model. Across four domains, it achieves a 60.8% attack success rate. On CNN/DailyMail, it reaches 70.2%.
In short: text can be made to look like it was written by one model even when it was originally generated by another.
Evading a source is not reaching a target
Most work on fooling these detectors only asks the text to stop looking like its real source — any other answer counts as a win. Landing on the one model you picked is much harder.
The two goals sound similar and score very differently. Planner TST and ForgePrint-4B are 10.5 points apart at simply escaping the source, but 27.7 points apart at reaching the intended target (33.1% vs. 60.8%).
How ForgePrint works
No corpus of paired rewrites exists — one model’s text rewritten to look like another’s — so ForgePrint builds its own. For each source→target pair it assembles editing operators (drop a habit of the source, adopt one of the target’s, reshape the sentences) and keeps only those that measurably move a surrogate classifier, a stand-in it trains itself, toward the target. A large Teacher applies them, generates several candidate rewrites and selects the best. A 4B Student is distilled from those, then refined on its own output, ranked by factual accuracy first and target resemblance second, so it cannot succeed by quietly altering what the text says.
Only the Student is deployed. Given the text and the two model names it rewrites in a single pass — no operator bank, no surrogate, no access to the original article.
Results
The attack never sees what grades it. It is developed against its own surrogate, trained on one half of the data. Every number below comes from four held-out classifiers — RoBERTa, DeBERTa, GPT-2 and TF-IDF, trained on the other half, which shares no documents with the first — and the attack never queries them.
The 4B Student beats the 26B Teacher it learned from (70.2% vs. 54.1% on news), because its last training stage lets it practise on its own rewrites instead of only copying the Teacher. The ordering holds across all four kinds of text — news, research papers, dialogue and how-to guides — and all four classifiers.
The gain is not a loss of faithfulness
A rewrite that quietly altered the content would score well and mean nothing. Compared only against rewrites of matched faithfulness, the lead over the strongest published method holds at 27.5 points under AlignScore and 28.3 under MiniCheck.
▸Table 2 — the faithfulness-adjusted numbers
Which target you aim at matters more than where you start
Claude is the hardest model to imitate, reached 39.7% of the time on average against 64.1–70.3% for the others — and how hard it is depends on what is being written: aiming at Claude works 5% of the time on research-paper summaries and 70% on news. Which model you aim at accounts for 82.6% of the difference between one rewriting job and another; which model you start from accounts for the other 17.4%.
From an open model to a chosen commercial one
Here an open model, Gemma, writes the news summaries, and the rewrite aims them at one named commercial model. Left alone, a Gemma summary is almost never mistaken for Claude (0.1%) or Grok (0.5%). After rewriting, it passes for the chosen model between 49.6% and 81.5% of the time.
Here the classifier can also answer “Gemma”, so a rewrite has to leave its real source and reach the model it was aimed at. Not on the same scale as Table 1.
The Student beats its Teacher on three of the four targets, most sharply on GPT (64.6% vs. 17.3%). A Student built on a different backbone, Qwen3.5-9B, reaches 66.7% on average, so this is not specific to the Gemma family.
Case study
cnndm_CorpusB:22:Grok_to_GeminiManchester United is preparing for its upcoming UEFA U19 Youth League campaign by conducting trials of promising young players. Specifically, the club will assess 18-year-old defender Luke Tingey and 18-year-old midfielder Kyran Wiltshire. Both players are members of the MK Dons U18 squad that recently won the Youth Alliance South Cup. This trial period will involve training sessions at Carrington.
Turns two long sentences into four short ones with one point each, opens with a framing sentence, and drops the free-kick distance and the adjective “lively”. Three of the four evaluators now name Gemini.
The Student is not copying Gemini’s words but the way it arranges them: two long sentences become four short ones making one point each, a framing sentence opens the piece, and incidental detail is dropped. The strongest published method instead reaches for exclamations and evaluative adjectives — and leaves behind a stray bracket from its own prompt template.
Takeaways
- A classifier can be right most of the time and still be wrong about where a particular text began. Our held-out classifiers reach 85.9% accuracy on unchanged summaries, but ForgePrint achieves a 70.2% attack success rate on CNN/DailyMail while preserving their meaning.
- Simply making a source label disappear is not the same as choosing a new one. Looking only at evasion hides the gap: a 10.5-point difference in evasion becomes a 27.7-point difference when the goal is to redirect attribution to a specific model.
- This exposes a real-world substitution threat: a provider could serve a small open-source model while making its outputs look like they came from a commercial API. Using a 4B rewriter, we move open-model outputs toward a chosen commercial-model label with a 68.3% attack success rate.
Ethics and release
Targeted rewriting is dual-use — the capability that measures the weakness is the one that exploits it — so we keep the experiments contained: everything runs offline on generated summaries and offline classifiers, we query commercial APIs only to produce summaries, and we never test impersonation against a deployed service.
Our code and evaluation suite will be released at github.com/HaohanYuan01/ForgePrint. The trained rewriter is available on request under the terms of the Ethics Statement.
BibTeX
@misc{yuan2026forgeprint, title = {Forging LLM Authorship Fingerprints with Targeted Rewriting}, author = {Yuan, Haohan and Chen, Simin and Niu, Xi and Guo, Hanqing and Xu, Depeng and Zhang, Haopeng}, year = {2026}, eprint = {2609.38831}, archivePrefix = {arXiv}, primaryClass = {cs.CL}}