PageSourceSearch

https://jerick-1380.github.io/publications/

html jerick-1380.github.io collected 2026-10-03 09:22:50 UTC 14,350 bytes, 206 lines download raw bytes

1<!doctype html>
2<html lang="en">
3<head>
4  <meta charset="utf-8">
5  <meta name="viewport" content="width=device-width, initial-scale=1">
6  <title>Research — Jerick Shi</title>
7  <meta name="description" content="Publications by Jerick Shi on multi-agent LLM systems, deception, and AI safety.">
8  <meta property="og:type" content="website">
9  <meta property="og:site_name" content="Jerick Shi">
10  <meta property="og:url" content="https://jerick-1380.github.io/publications/">
11  <meta property="og:title" content="Research">
12  <meta property="og:description" content="Publications by Jerick Shi on multi-agent LLM systems, deception, and AI safety.">
13  <meta property="og:image" content="https://jerick-1380.github.io/assets/img/social/og-default.png">
14  <meta property="og:image:width" content="1200">
15  <meta property="og:image:height" content="627">
16  <meta name="twitter:card" content="summary_large_image">
17  <link rel="icon" href="/assets/img/brain.jpg">
18  <link rel="preconnect" href="https://fonts.googleapis.com">
19  <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
20  <link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600&family=JetBrains+Mono:wght@400;700&family=Space+Grotesk:wght@500;600;700&display=swap" rel="stylesheet">
21  <link rel="stylesheet" href="/assets/css/main.css?v=202609271451">
22</head>
23<body>
24  <div id="progress"></div>
25  <canvas id="starfield"></canvas>
26  <div class="bg-glow"></div>
27
28  <nav class="nav">
29    <div class="nav-inner">
30      <a class="logo" href="/"><b>jerick</b>@shi:~$<span class="cursor">▊</span></a>
31      <button class="nav-burger" aria-label="Menu">☰</button>
32      <div class="nav-links">
33        <a href="/">About</a>
34        <a href="/publications/">Research</a>
35        <a href="/projects/">Projects</a>
36        <a href="/teaching/">Teaching</a>
37        <a href="/blog/">Blog</a>
38        <a href="/books/">Books</a>
39        <a href="/hobbies/">Hobbies</a>
40        <a href="/cv/">CV</a>
41        <button class="cmdk-btn" aria-label="Open command palette">⌘K</button>
42      </div>
43    </div>
44  </nav>
45
46  <header class="page-head wrap">
47    <div class="eyebrow"><span class="idx">//</span> research_index</div>
48    <h1 data-scramble>Publications</h1>
49    <p class="lead">Deception, communication, and trust in multi-agent LLM systems. Also on <a href="https://scholar.google.com/citations?user=6wj2mTQAAAAJ" target="_blank" rel="noopener">Google Scholar</a>.</p>
50  </header>
51
52  <main class="wrap" style="padding-bottom:3rem;">
53
54    <!-- Research map -->
55    <section id="map" class="reveal" style="margin-bottom:3.5rem;">
56      <div class="eyebrow"><span class="idx">//</span> research_graph</div>
57      <p class="muted" style="max-width:640px; margin-bottom:1.2rem; font-size:0.95rem;">Every paper, the ideas it touches, and the people behind it, as a live force-directed graph. Drag nodes, hover to trace connections, click a paper to open it. Hovering a paper below lights it up here too.</p>
58      <div data-research-map aria-label="Interactive map of research papers, concepts, and collaborators"></div>
59    </section>
60
61    <!-- Thesis banner -->
62    <div class="card tilt reveal" data-node="thesis" style="margin-bottom:2.5rem; border-image: linear-gradient(120deg,#46e0ff,#9d6bff) 1;">
63      <span class="chip">Master's Thesis · CMU-CS-26-105</span>
64      <h3 style="margin-top:0.8rem;">The Structure of Deception: How LLM Agents Lie, Break Promises, and Exploit Trust in Multi-Agent Settings</h3>
65      <p class="muted" style="font-size:0.95rem;">A unified framework for measuring LLM deception across three interaction structures — one-shot games, repeated games, and open-ended resource simulations. Committee: Vincent Conitzer (chair), Aditi Raghunathan. Advisors: Vincent Conitzer, Zhijing Jin.</p>
66      <div class="pub-actions">
67        <a class="btn btn-sm btn-primary" href="/assets/pdf/masters_thesis.pdf" target="_blank">Thesis PDF</a>
68        <a class="btn btn-sm btn-ghost" href="https://youtu.be/Z3Q9AkriPxg" target="_blank" rel="noopener">Defense video</a>
69        <a class="btn btn-sm btn-ghost" href="/projects/masters-thesis/">Project page</a>
70      </div>
71    </div>
72
73    <h2 class="reveal" style="margin-top:3rem;">2026</h2>
74    <hr class="divider-glow" style="margin:1rem 0 2rem;">
75
76    <!-- When Agents Lie -->
77    <div class="card reveal" data-node="lie" style="margin-bottom:1.6rem; border-image: linear-gradient(120deg,#ffd479,#9d6bff) 1;">
78      <span class="chip" style="color:#ffd479; border-color:rgba(255,212,121,0.35); background:rgba(255,212,121,0.07);">★ BEST PAPER AWARD</span>
79      <h3 style="margin-top:0.8rem;">When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games</h3>
80      <p class="authors"><b>Jerick Shi</b>, Terry Jingchen Zhang, Bernhard Schölkopf, Vincent Conitzer, Zhijing Jin</p>
81      <p class="venue">ICML 2026 · NEW FRONTIERS IN GAME-THEORETIC LEARNING (NEXT-GAME) WORKSHOP · PRESENTED JULY 11, SEOUL</p>
82      <div class="pub-actions">
83        <a class="btn btn-sm btn-primary" href="https://openreview.net/pdf?id=v8nYIkYjY0" target="_blank" rel="noopener">Paper (PDF)</a>
84        <a class="btn btn-sm btn-ghost" href="https://openreview.net/forum?id=v8nYIkYjY0" target="_blank" rel="noopener">OpenReview</a>
85        <button class="btn btn-sm btn-ghost" data-copy="bib-lie">Copy BibTeX</button>
86      </div>
87      <pre id="bib-lie" hidden>@inproceedings{shi2026when,
88  title={When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games},
89  author={Jerick Shi and Terry Jingchen Zhang and Bernhard Sch{\"o}lkopf and Vincent Conitzer and Zhijing Jin},
90  booktitle={ICML 2026 Workshop on New Frontiers in Game-Theoretic Learning (NExT-Game)},
91  year={2026},
92  note={Best Paper Award},
93  url={https://openreview.net/forum?id=v8nYIkYjY0}
94}</pre>
95    </div>
96
97    <!-- Strategic Silence -->
98    <div class="card reveal" data-node="silence" style="margin-bottom:1.6rem;">
99      <h3>What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs</h3>
100      <p class="authors"><b>Jerick Shi</b>, Terry Jingchen Zhang, Vincent Conitzer, Zhijing Jin</p>
101      <p class="venue">ICML 2026 · WORKSHOP ON FAILURE MODES OF AGENTIC AI</p>
102      <div class="pub-actions">
103        <a class="btn btn-sm btn-primary" href="https://openreview.net/pdf?id=ZOdCsExYgi" target="_blank" rel="noopener">Paper (PDF)</a>
104        <a class="btn btn-sm btn-ghost" href="https://openreview.net/forum?id=ZOdCsExYgi" target="_blank" rel="noopener">OpenReview</a>
105        <button class="btn btn-sm btn-ghost" data-copy="bib-silence">Copy BibTeX</button>
106      </div>
107      <pre id="bib-silence" hidden>@inproceedings{shi2026silence,
108  title={What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs},
109  author={Jerick Shi and Terry Jingchen Zhang and Vincent Conitzer and Zhijing Jin},
110  booktitle={ICML 2026 Workshop on Failure Modes of Agentic AI},
111  year={2026},
112  url={https://openreview.net/forum?id=ZOdCsExYgi}
113}</pre>
114    </div>
115
116    <!-- Cheap Talk -->
117    <div class="card reveal pub" data-node="cheap" style="margin-bottom:1.6rem;">
118      <img src="/assets/img/publication_preview/cheap_talk.png" alt="Cheap Talk paper preview">
119      <div>
120        <h3>Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest</h3>
121        <p class="authors"><b>Jerick Shi</b>, Terry Jingcheng Zhang, Zhijing Jin, Vincent Conitzer</p>
122        <p class="venue">ICLR 2026 · AI FOR MECHANISM DESIGN &amp; STRATEGIC DECISION MAKING WORKSHOP · ARXIV:2604.04782</p>
123        <div class="pub-actions">
124          <a class="btn btn-sm btn-primary" href="https://arxiv.org/abs/2604.04782" target="_blank" rel="noopener">arXiv</a>
125          <button class="btn btn-sm btn-ghost" data-toggle="abs-cheap">Abstract</button>
126          <button class="btn btn-sm btn-ghost" data-copy="bib-cheap">Copy BibTeX</button>
127        </div>
128        <div class="abstract" id="abs-cheap">Large language models are increasingly deployed as autonomous agents in multi-agent settings where they communicate intentions and take consequential actions with limited human oversight. A critical safety question is whether agents that publicly commit to actions break those commitments when they can privately deviate, and what the consequences are for both themselves and the collective. We study deception as a deviation from a publicly announced action in one-shot normal-form games, classifying each deviation by its effect on individual payoff and collective welfare into four categories: strategic, selfish, altruistic, and sabotaging. By exhaustively enumerating announcement profiles across six canonical games and nine frontier models, we identify all opportunities for each deviation type and measure how often agents exploit them. Across all settings, agents deviate from commitments in approximately 56.6% of scenarios, but the character of deception varies substantially across models even at similar overall rates. Most critically, for the majority of the models, commitment-breaking occurs without metacognitive awareness as measured by LLM-judged reasoning traces, with agents optimizing payoffs without recognizing that they are breaking commitments.</div>
129        <pre id="bib-cheap" hidden>@article{deception26,
130  title={Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest},
131  author={Jerick Shi and Terry Jingcheng Zhang and Zhijing Jin and Vincent Conitzer},
132  journal={arXiv preprint arXiv:2604.04782},
133  year={2026},
134  url={https://arxiv.org/abs/2604.04782}
135}</pre>
136      </div>
137    </div>
138
139    <!-- Taxonomy -->
140    <div class="card reveal pub" data-node="tax" style="margin-bottom:1.6rem;">
141      <img src="/assets/img/publication_preview/taxonomy.png" alt="Deception taxonomy paper preview">
142      <div>
143        <h3>From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception</h3>
144        <p class="authors"><b>Jerick Shi</b>, Terry Jingcheng Zhang, Zhijing Jin, Vincent Conitzer</p>
145        <p class="venue">ICLR 2026 · AGENTS IN THE WILD: SAFETY, SECURITY &amp; BEYOND WORKSHOP · ARXIV:2604.04788</p>
146        <div class="pub-actions">
147          <a class="btn btn-sm btn-primary" href="https://arxiv.org/abs/2604.04788" target="_blank" rel="noopener">arXiv</a>
148          <button class="btn btn-sm btn-ghost" data-toggle="abs-tax">Abstract</button>
149          <button class="btn btn-sm btn-ghost" data-copy="bib-tax">Copy BibTeX</button>
150        </div>
151        <div class="abstract" id="abs-tax">Large language models produce outputs that systematically mislead users, from hallucinated facts and fabricated citations to sycophantic agreement and strategic deception of evaluators. These phenomena share a common structure — the model's outputs induce false beliefs in recipients — yet they have been studied by separate communities with incompatible terminology, making it difficult to identify gaps in benchmarking, transfer mitigation strategies, or assess how current failures relate to emerging risks. We propose a unified taxonomy organized along three dimensions: behavioral versus strategic deception (whether misleading outputs are training artifacts or instrumentally selected), objects of misrepresentation (what is misrepresented, across seven categories from factual claims to 
151stated objectives), and mechanisms (commission, omission, or pragmatic distortion). Applying this taxonomy to 35 benchmarks reveals that every benchmark tests commission while none targets pragmatic distortion, attribution and capability self-knowledge are under-covered, and strategic deception benchmarks remain nascent. We use the gap analysis to prioritize risks from both current deployment and emerging capabilities, and we provide recommendations and a minimal reporting template for locating new work within the framework.</div>
152        <pre id="bib-tax" hidden>@article{survey26,
153  title={From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception},
154  author={Jerick Shi and Terry Jingcheng Zhang and Zhijing Jin and Vincent Conitzer},
155  journal={arXiv preprint arXiv:2604.04788},
156  year={2026},
157  url={https://arxiv.org/abs/2604.04788}
158}</pre>
159      </div>
160    </div>
161
162    <h2 class="reveal" style="margin-top:3rem;">2025</h2>
163    <hr class="divider-glow" style="margin:1rem 0 2rem;">
164
165    <!-- Market -->
166    <div class="card reveal pub" data-node="market" style="margin-bottom:1.6rem;">
167      <img src="/assets/img/publication_preview/market.png" alt="Market communication paper preview">
168      <div>
169        <h3>Market-Dependent Communication in Multi-Agent Alpha Generation</h3>
170        <p class="authors"><b>Jerick Shi</b>, Burton Hollifield</p>
171        <p class="venue">NEURIPS 2025 · GENAI IN FINANCE WORKSHOP · ARXIV:2511.13614</p>
172        <div class="pub-actions">
173          <a class="btn btn-sm btn-primary" href="https://doi.org/10.48550/arXiv.2511.13614" target="_blank" rel="noopener">arXiv</a>
174          <button class="btn btn-sm btn-ghost" data-copy="bib-market">Copy BibTeX</button>
175        </div>
176        <pre id="bib-market" hidden>@article{shi2025market,
177  title={Market-Dependent Communication in Multi-Agent Alpha Generation},
178  author={Jerick Shi and Burton Hollifield},
179  journal={CoRR},
180  volume={abs/2511.13614},
181  year={2025},
182  url={https://doi.org/10.48550/arXiv.2511.13614}
183}</pre>
184      </div>
185    </div>
186
187  </main>
188
189  <footer>
190    <div class="wrap foot-grid">
191      <div class="foot-links">
192        <a href="mailto:[email protected]">Email</a>
193        <a href="https://github.com/Jerick-1380" target="_blank" rel="noopener">GitHub</a>
194        <a href="https://www.linkedin.com/in/jerick-shi-293773216" target="_blank" rel="noopener">LinkedIn</a>
195        <a href="https://scholar.google.com/citations?user=6wj2mTQAAAAJ" target="_blank" rel="noopener">Scholar</a>
196        <a href="https://www.youtube.com/@DummyR18" target="_blank" rel="noopener">YouTube</a>
197        <a href="https://flickr.com/photos/203834484@N07/" target="_blank" rel="noopener">Flickr</a>
198        <a href="/books/">Bookshelf</a>
199      </div>
200      <div class="foot-note">© 2026 JERICK SHI · PRESS ⌘K TO NAVIGATE</div>
201    </div>
202  </footer>
203
204  
204<script src="/assets/js/main.js?v=202609271451"></script>
204
205</body>
206</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.