1<!doctype html> 2<html lang="en"> 3<head> 4 <meta charset="utf-8"> 5 <title>Leveraging Human Expertise for LLM-Assisted Dialogue Character Extraction and Attribution in Classic Chinese Novel</title> 6 <meta name="viewport" content="width=device-width, initial-scale=1.0"> 7 <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/dist/reveal.css"> 8 <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/dist/theme/black.css" id="theme"> 9 <style> 10 .two-column { 11 display: flex; 12 gap: 2rem; 13 } 14 .two-column > div { 15 flex: 1; 16 } 17 .image-placeholder { 18 background: #f0f0f0; 19 border: 2px dashed #ccc; 20 padding: 2rem; 21 text-align: center; 22 border-radius: 8px; 23 margin: 1rem 0; 24 } 25 .small-text { 26 font-size: 0.8em; 27 } 28 .ssmall-text { 29 font-size: 0.5em; 30 } 31 .sssmall-text { 32 white-space: pre-line; 33 font-size: 0.5em; 34 text-align: left; 35 } 36 .slide-image { 37 max-width: 90%; 38 height: auto; 39 border-radius: 8px; 40 box-shadow: 0 4px 8px rgba(0,0,0,0.1); 41 margin: 1rem auto; 42 display: block; 43 } 44 .title-image { 45 max-width: 80%; 46 height: auto; 47 border-radius: 8px; 48 box-shadow: 0 4px 8px rgba(0,0,0,0.1); 49 margin: 2rem auto; 50 display: block; 51 } 52 .snap_containter img { 53 height: 360px; 54 width: auto; 55 margin-right: 1%; 56 vertical-align: middle; 57 object-fit: cover; 58 } 59 </style> 60</head> 61<body> 62 <div class="reveal"> 63 <div class="slides"> 64 <!-- Title Slide --> 65 <section> 66 <h1 style="font-size: 1em;">Leveraging Human Expertise for LLM-Assisted Dialogue Character Extraction and Attribution in Classic Chinese Novel</h1> 67 <h3 style="font-size: .5em;">Yutong Yang, Shanghai Jiao Tong University</h3> 68 <!-- å端çé¢å±ç¤º --> 69 <img src="figure/frontend_all.png" alt="Frontend Interface" class="title-image"> 70 </section> 71 72 <!-- 1. Introduction --> 73 <section> 74 <section> 75 <h2>1. Introduction</h2> 76 <div class="two-column"> 77 <div> 78 <p class="small-text">Classic Chinese Novels - Using Distance Reading to Understand Complex Narratives</p> 79 <p class="ssmall-text">· Revealing new insights into character relationships and narrative roles...</p> 80 </div> 81 <div> 82 <p class="small-text">Social Network Analysis - Requiring Accurate Data Extracted from the Novels</p> 83 <p class="ssmall-text">· Challenges in classic Chinese novels</p> 84 <p class="ssmall-text">· Challenges in data extraction</p> 85 </div> 86 </div> 87 </section> 88 89 <section> 90 <h4>Challenges in Classic Chinese Novels</h4> 91 <div class="two-column"> 92 <div style="flex: 2;"> 93 <ul class="small-text"> 94 <li>Complex character networks</li> 95 <li>Rich contexts</li> 96 <li>Multiple aliases per character</li> 97 </ul> 98 </div> 99 <div style="flex: 1;"> 100 <!-- éè¦å¾çï¼çº¢æ¥¼æ¢¦äººç©å ³ç³»ç½ç»å¾ --> 101 <img src="figure/book.png" alt="红楼梦人ç©å ³ç³»ç½ç»å¾" class="slide-image" style="max-width: 80%;"> 102 <p class="sssmall-text">Dream of the Red Chamber</p> 103 </div> 104 105 </div> 106 <img src="figure/daguanyuan.jpg" alt="红楼梦人ç©å ³ç³»ç½ç»å¾" style="max-width: 30%;"> 107 <img src="figure/daguanyuan2.jpg" alt="红楼梦人ç©å ³ç³»ç½ç»å¾" style="max-width: 30%;"> 108 <img src="figure/daguanyuan4.jpg" alt="红楼梦人ç©å ³ç³»ç½ç»å¾"style="max-width: 30%;"> 109 </section> 110 111 <section> 112 <h4>Challenges in Data Extraction</h4> 113 <div class="two-column"> 114 <div style="flex: 2;"> 115 <ul class="small-text"> 116 <li>Manual extraction of characters from classic Chinese novels requires significant human effort and is highly time-consuming.</li> 117 <li>The richness of context in novels and the complexity of character references present challenges for fully automated extraction methods.</li> 118 </ul> 119 </div> 120 121 </div> 122 </section> 123 </section> 124 125 <!-- 2. Framework Overview --> 126 <section> 127 <section> 128 <h2>2. Framework Overview</h2> 129 <div class="two-column"> 130 <div style="flex: 0.5;"> 131 <div style="margin-bottom: 1rem;"> 132 <h5>Back-end Processing</h5> 133 <ul class="ssmall-text"> 134 <li>Data extraction</li> 135 <li>LLM processing</li> 136 <li>Character identification</li> 137 </ul> 138 </div> 139 <div> 140 <h5>Front-end Interface</h5> 141 <ul class="ssmall-text"> 142 <li>Interactive annotation</li> 143 <li>Visualization</li> 144 <li>Manual refinement</li> 145 </ul> 146 </div> 147 </div> 148 <div style="flex: 1;"> 149 <!-- éè¦å¾çï¼æ´ä½æ¡æ¶æ¶æå¾ --> 150 <img src="figure/pipeline.png" alt="æ´ä½æ¡æ¶æ¶æå¾" class="slide-image"> 151 </div> 152 </div> 153 </section> 154 </section> 155 156 <!-- 3. Back-end Data Processing --> 157 <section> 158 <section> 159 <h2>3. Back-end Data Processing</h2> 160 </section> 161 162 <section> 163 <h3>3.1 Upload and Segmentation</h3> 164 <div class="two-column"> 165 <div> 166 <h4>Text Upload</h4> 167 <p class="small-text">Original novel text upload</p> 168 </div> 169 <div> 170 <h4>Rule-based Segmentation</h4> 171 <p class="small-text">Break into dialogue units with context</p> 172 </div> 173 </div> 174 <div class="snap_containter"> 175 <img src="figure/backend_uoload.png" alt="ææ¬åå²"> 176 <img src="figure/backend_cut.png" alt="ææ¬åå²"> 177 </div> 178 </section> 179 180 <section> 181 <h3>3.2 LLM-Based Extraction</h3> 182 <div class="two-column"> 183 <div> 184 <!-- <h4>Speaker/Listener Extraction</h4> --> 185 <p class="small-text">Speaker and Listener Extraction</p> 186 <p class="sssmall-text">prompt = f"\nQ: I will give you a dialogue sentence and a passage of context. Please repeat the dialogue sentence, then based on the context, identify the speaker, the primary listener(s), and the secondary listener(s) in the specified dialogue sentence. 187 Dialogue sentence: {talk}. 188 Context: {context}. 189 Please provide your answer in the format: 190 'Dialogue sentence: [dialogue], Speaker: [speaker], Primary Listener(s): [primary listener(s)], Secondary Listener(s): [secondary listener(s)]'. 191 Note: 192 1. Use commas (",") to separate the dialogue sentence, speaker, primary listener(s), and secondary listener(s); 193 2. If there are multiple speakers or multiple primary/secondary listeners, separate them using "ã"; 194 3. Do not insert any line breaks in your answer; 195 4. Resolve pronoun references carefully and avoid vague references like "you", "I", "he", etc.; 196 5. Do not include any explanation or analysis in your answer â treat it like a fill-in-the-blank question. The more concise, the better; 197 6. If the speaker, primary listener(s), or secondary listener(s) cannot be identified, respond with 'None'. \nA:" 198 </p> 199 <p class="small-text">⢠Add to Database</p> 200 </div> 201 </div> 202 </section> 203 204 <section> 205 <h3>3.3 LLM-Assisted Attribution</h3> 206 <div class="two-column"> 207 <div style="flex: 2;"> 208 <!-- <h4>Name Chain Database</h4> --> 209 <p class="small-text">Resolve Co-references</p> 210 <p class="sssmall-text">prompt = 211 f"\nQ:This is a fill-in-the-blank question with an answer of 0 or 1. Please determine: In the book '{chinese_book_name}', are '{entity}' and '{main_entity}' the same character? 212 If yes, return 1; if not, return 0. Do not consider literary implications. No explanation or analysis is needed." 213 \nA:" 214 </p> 215 <p class="small-text">⢠Add to Database</p> 216 <!-- <img src="figure/backend_entitychain.png" alt="人称é¾" style="height: 25%;"> --> 217 </div> 218 <!-- <div style="flex: 1;"> 219 <h4>Dialogue Network</h4> 220 <p class="small-text">Visualize interactions</p> 221 </div> --> 222 </div> 223 <!-- <div class="image-placeholder"> 224 [éè¦å¾çï¼å¯¹è¯ç½ç»çæå¾] 225 </div> --> 226 </section> 227 </section> 228 229 <!-- 4. Front-end Annotation Interface --> 230 <section> 231 <section> 232 <h2>4. Front-end Annotation Interface</h2> 233 <img src="figure/frontend_all.png" alt="Frontend Interface" > 234 </section> 235 236 <section> 237 <h3>4.1 Worktable</h3> 238 <div class="two-column"> 239 <div style="flex: 0.5;"> 240 <div style="margin-bottom: 1rem;"> 241 <p class="small-text">Upload & Cut Dialogues</p> 242 </div> 243 </div> 244 <!-- å端ç颿ªå¾ --> 245 <div style="flex: 1;"> 246 <img src="figure/frontend_upload.jpg" alt="Frontend Interface" class="slide-image"> 247 </div> 248 </div> 249 </section> 250 251 <section> 252 <h3>4.2 Interactive Annotation</h3> 253 <div class="two-column"> 254 <div style="flex: 0.5;"> 255 <div> 256 <p class="small-text">Refine annotations: click to switch roles</p> 257 </div> 258 </div> 259 <!-- å端ç颿ªå¾ --> 260 <!-- <div style="flex: 1;"> 261 <img src="figure/frontend_annotation.png" alt="Frontend Interface" class="slide-image"> 262 </div> --> 263 <div style="flex: 1;"> 264 <video src="figure/annotation.mp4" controls autoplay muted loop playsinline class="slide-image"> 265 Your browser does not support the video tag. 266 </video> 267 </div> 268 </div> 269 </section> 270 271 <!-- <section> 272 <video 273 src="figure/annotation.mp4" 274 controls 275 autoplay 276 loop 277 muted 278 playsinline 279 style="max-width: 100%; height: auto; "> 280 </video> 281 </section> --> 282 283 <section> 284 <h3>4.3 Visualization Features</h3> 285 <div class="two-column"> 286 <div style="flex: 1;"> 287 <h5>Data visualization design</h5> 288 <ul class="ssmall-text"> 289 <li>Large nodes (green) = main name</li> 290 <li>Small nodes (yellow) = aliases</li> 291 <li>Links between large nodes = relations (measured with numbers of conversation)</li> 292 <li>Chapter-specific networks: generating dialogue networks for all character relationships up to the current chapter, reveal
292ing narrative progression.</li> 293 </ul> 294 </div> 295 <div style="flex: 1;"> 296 <video src="figure/network.mp4" controls autoplay muted loop playsinline class="slide-image"> 297 Your browser does not support the video tag. 298 </video> 299 </div> 300 </div> 301 </section> 302 303 <section> 304 <h3>4.3 Visualization Features</h3> 305 <div class="two-column"> 306 <div style="flex: 0.5;"> 307 <h5>Manual Disambiguation</h5> 308 <p class="ssmall-text">Correct extraction errors</p> 309 </div> 310 <div style="flex: 1;"> 311 <video src="figure/disambiguation.mp4" controls autoplay muted loop playsinline class="slide-image"> 312 Your browser does not support the video tag. 313 </video> 314 </div> 315 </div> 316 </section> 317 </section> 318 319 <!-- 5. Case Studies --> 320 <section> 321 <section> 322 <h2>5. Simple Insights from the Data Visualization</h2> 323 </section> 324 325 <!-- <section> 326 <h3>5.1 Dream of the Red Chamber</h3> 327 <div class="two-column"> 328 <div> 329 <h4>Initial Extraction</h4> 330 <p class="small-text">LLM-based character identification</p> 331 </div> 332 <div> 333 <h4>Manual Refinement</h4> 334 <p class="small-text">Human expertise correction</p> 335 </div> 336 </div> 337 <img src="figure/frontend_annotation.png" alt="case" style="width: 70%;"> 338 </section> --> 339 340 <section> 341 <!-- <h3>Data Visualization</h3> --> 342 <div class="two-column"> 343 <div> 344 <h4>Key Characters</h4> 345 <p class="small-text">Important intermediary characters identified</p> 346 </div> 347 <div> 348 <h4>Interaction Patterns</h4> 349 <p class="small-text">Social dynamics revealed</p> 350 </div> 351 </div> 352 <img src="figure/chapters.png" alt="case" style="width: 70%;"> 353 354 </section> 355 <section> 356 <h4>Network of the first chapter</h4> 357 <img src="figure/first_chapter_net.png" alt="case" style="width: 100%;"> 358 </section> 359 <section> 360 <h4>Network of the second chapter</h4> 361 <img src="figure/second_chapter_net.png" alt="case" style="width: 100%;"> 362 </section> 363 <section> 364 <h4>Network of the third chapter</h4> 365 <img src="figure/third_chapter_net.png" alt="case" style="width: 100%;"> 366 </section> 367 <section> 368 <h4>Network of the forth chapter</h4> 369 <img src="figure/forth_chapter_net.png" alt="case" style="width: 100%;"> 370 </section> 371 </section> 372 373 <!-- 6. Conclusion --> 374 <section> 375 <section> 376 <h2>6. Conclusion</h2> 377 </section> 378 379 <section> 380 <h3>6.1 Summary</h3> 381 <div class="two-column"> 382 <div> 383 <p class="small-text">AI + Human</p> 384 </div> 385 <img src="figure/pipeline.png" alt="case" style="width: 70%;"> 386 </div> 387 </section> 388 389 <section> 390 <h3>6.2 TODOs</h3> 391 <div class="two-column"> 392 <div> 393 <h4>Improving Algorithm Efficiency</h4> 394 <p class="small-text">Improve the speed of data processing, thus generate the result faster.</p> 395 </div> 396 <div> 397 <h4>Incorporating SNA Algorithm</h4> 398 <p class="small-text">Use social network analysis algorithm to build a final network, which can directly serve for the literary analysis.</p> 399 </div> 400 </div> 401 </section> 402 </section> 403 404 <!-- Thank You --> 405 <section> 406 <h2>Thank You!</h2> 407 <p class="small-text" style="font-style: italic;">[email protected]</p> 408 <p class="small-text" style="font-style: italic;">https://yutong-yang.github.io/</p> 409 </section> 410 </div> 411 </div> 412
412<script src="https://cdn.jsdelivr.net/npm/[email protected]/dist/reveal.js"></script>
412 413
413<script> 414 Reveal.initialize({ 415 hash: true, 416 transition: 'slide' 417 }); 418 </script>
418 419</body> 420</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.