PageSourceSearch

https://yutong-yang.github.io/dh2025_pre_html/index.html

html yutong-yang.github.io collected 2026-10-03 10:13:54 UTC 16,231 bytes, 420 lines download raw bytes

1<!doctype html>
2<html lang="en">
3<head>
4  <meta charset="utf-8">
5  <title>Leveraging Human Expertise for LLM-Assisted Dialogue Character Extraction and Attribution in Classic Chinese Novel</title>
6  <meta name="viewport" content="width=device-width, initial-scale=1.0">
7  <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/dist/reveal.css">
8  <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/dist/theme/black.css" id="theme">
9  <style>
10    .two-column {
11      display: flex;
12      gap: 2rem;
13    }
14    .two-column > div {
15      flex: 1;
16    }
17    .image-placeholder {
18      background: #f0f0f0;
19      border: 2px dashed #ccc;
20      padding: 2rem;
21      text-align: center;
22      border-radius: 8px;
23      margin: 1rem 0;
24    }
25    .small-text {
26      font-size: 0.8em;
27    }
28    .ssmall-text {
29      font-size: 0.5em;
30    }
31    .sssmall-text {
32      white-space: pre-line;
33      font-size: 0.5em;
34      text-align: left;
35    }
36    .slide-image {
37      max-width: 90%;
38      height: auto;
39      border-radius: 8px;
40      box-shadow: 0 4px 8px rgba(0,0,0,0.1);
41      margin: 1rem auto;
42      display: block;
43    }
44    .title-image {
45      max-width: 80%;
46      height: auto;
47      border-radius: 8px;
48      box-shadow: 0 4px 8px rgba(0,0,0,0.1);
49      margin: 2rem auto;
50      display: block;
51    }
52    .snap_containter img {
53      height: 360px;
54      width: auto;
55      margin-right: 1%;
56      vertical-align: middle;
57      object-fit: cover;
58    }
59  </style>
60</head>
61<body>
62  <div class="reveal">
63    <div class="slides">
64      <!-- Title Slide -->
65      <section>
66        <h1 style="font-size: 1em;">Leveraging Human Expertise for LLM-Assisted Dialogue Character Extraction and Attribution in Classic Chinese Novel</h1>
67        <h3 style="font-size: .5em;">Yutong Yang, Shanghai Jiao Tong University</h3>
68        <!-- 前端界面展示 -->
69        <img src="figure/frontend_all.png" alt="Frontend Interface" class="title-image">
70      </section>
71
72      <!-- 1. Introduction -->
73      <section>
74        <section>
75          <h2>1. Introduction</h2>
76          <div class="two-column">
77            <div>
78              <p class="small-text">Classic Chinese Novels - Using Distance Reading to Understand Complex Narratives</p>
79              <p class="ssmall-text">· Revealing new insights into character relationships and narrative roles...</p>
80            </div>
81            <div>
82              <p class="small-text">Social Network Analysis - Requiring Accurate Data Extracted from the Novels</p>
83              <p class="ssmall-text">· Challenges in classic Chinese novels</p>
84              <p class="ssmall-text">· Challenges in data extraction</p>
85            </div>
86          </div>
87        </section>
88        
89        <section>
90          <h4>Challenges in Classic Chinese Novels</h4>
91          <div class="two-column">
92            <div style="flex: 2;">
93              <ul class="small-text">
94                <li>Complex character networks</li>
95                <li>Rich contexts</li>
96                <li>Multiple aliases per character</li>
97              </ul>
98            </div>
99            <div style="flex: 1;">
100              <!-- 需要图片:红楼梦人物关系网络图 -->
101              <img src="figure/book.png" alt="红楼梦人物关系网络图" class="slide-image" style="max-width: 80%;">
102              <p class="sssmall-text">Dream of the Red Chamber</p>
103            </div>
104            
105          </div>
106          <img src="figure/daguanyuan.jpg" alt="红楼梦人物关系网络图" style="max-width: 30%;">
107          <img src="figure/daguanyuan2.jpg" alt="红楼梦人物关系网络图" style="max-width: 30%;">
108          <img src="figure/daguanyuan4.jpg" alt="红楼梦人物关系网络图"style="max-width: 30%;">
109        </section>
110
111      <section>
112        <h4>Challenges in Data Extraction</h4>
113        <div class="two-column">
114          <div style="flex: 2;">
115            <ul class="small-text">
116              <li>Manual extraction of characters from classic Chinese novels requires significant human effort and is highly time-consuming.</li>
117              <li>The richness of context in novels and the complexity of character references present challenges for fully automated extraction methods.</li>
118            </ul>
119          </div>
120          
121        </div>
122      </section>
123    </section>
124
125      <!-- 2. Framework Overview -->
126      <section>
127        <section>
128          <h2>2. Framework Overview</h2>
129          <div class="two-column">
130            <div style="flex: 0.5;">
131              <div style="margin-bottom: 1rem;">
132                <h5>Back-end Processing</h5>
133                <ul class="ssmall-text">
134                  <li>Data extraction</li>
135                  <li>LLM processing</li>
136                  <li>Character identification</li>
137                </ul>
138              </div>
139              <div>
140                <h5>Front-end Interface</h5>
141                <ul class="ssmall-text">
142                  <li>Interactive annotation</li>
143                  <li>Visualization</li>
144                  <li>Manual refinement</li>
145                </ul>
146              </div>
147            </div>
148            <div style="flex: 1;">
149              <!-- 需要图片:整体框架架构图 -->
150              <img src="figure/pipeline.png" alt="整体框架架构图" class="slide-image">
151            </div>
152          </div>
153        </section>
154      </section>
155
156      <!-- 3. Back-end Data Processing -->
157      <section>
158        <section>
159          <h2>3. Back-end Data Processing</h2>
160        </section>
161        
162        <section>
163          <h3>3.1 Upload and Segmentation</h3>
164          <div class="two-column">
165            <div>
166              <h4>Text Upload</h4>
167              <p class="small-text">Original novel text upload</p>
168            </div>
169            <div>
170              <h4>Rule-based Segmentation</h4>
171              <p class="small-text">Break into dialogue units with context</p>
172            </div>
173          </div>
174          <div class="snap_containter">
175            <img src="figure/backend_uoload.png" alt="文本分割">
176            <img src="figure/backend_cut.png" alt="文本分割">
177          </div>
178        </section>
179
180        <section>
181          <h3>3.2 LLM-Based Extraction</h3>
182          <div class="two-column">
183            <div>
184              <!-- <h4>Speaker/Listener Extraction</h4> -->
185              <p class="small-text">Speaker and Listener Extraction</p>
186              <p class="sssmall-text">prompt = f"\nQ: I will give you a dialogue sentence and a passage of context. Please repeat the dialogue sentence, then based on the context, identify the speaker, the primary listener(s), and the secondary listener(s) in the specified dialogue sentence.  
187                Dialogue sentence: {talk}.  
188                Context: {context}.  
189                Please provide your answer in the format:  
190                'Dialogue sentence: [dialogue], Speaker: [speaker], Primary Listener(s): [primary listener(s)], Secondary Listener(s): [secondary listener(s)]'.  
191                Note:  
192                1. Use commas (",") to separate the dialogue sentence, speaker, primary listener(s), and secondary listener(s);  
193                2. If there are multiple speakers or multiple primary/secondary listeners, separate them using "、";  
194                3. Do not insert any line breaks in your answer;  
195                4. Resolve pronoun references carefully and avoid vague references like "you", "I", "he", etc.;  
196                5. Do not include any explanation or analysis in your answer — treat it like a fill-in-the-blank question. The more concise, the better;  
197                6. If the speaker, primary listener(s), or secondary listener(s) cannot be identified, respond with 'None'.  \nA:"
198              </p>
199              <p class="small-text">⇢ Add to Database</p>
200            </div>
201          </div>
202        </section>
203
204        <section>
205          <h3>3.3 LLM-Assisted Attribution</h3>
206          <div class="two-column">
207            <div style="flex: 2;">
208              <!-- <h4>Name Chain Database</h4> -->
209              <p class="small-text">Resolve Co-references</p>
210              <p class="sssmall-text">prompt = 
211                f"\nQ:This is a fill-in-the-blank question with an answer of 0 or 1. Please determine: In the book '{chinese_book_name}', are '{entity}' and '{main_entity}' the same character? 
212                If yes, return 1; if not, return 0. Do not consider literary implications. No explanation or analysis is needed."
213                \nA:"
214              </p>
215              <p class="small-text">⇢ Add to Database</p>
216              <!-- <img src="figure/backend_entitychain.png" alt="人称链" style="height: 25%;"> -->
217            </div>
218            <!-- <div style="flex: 1;">
219              <h4>Dialogue Network</h4>
220              <p class="small-text">Visualize interactions</p>
221            </div> -->
222          </div>
223          <!-- <div class="image-placeholder">
224            [需要图片:对话网络生成图]
225          </div> -->
226        </section>
227      </section>
228
229      <!-- 4. Front-end Annotation Interface -->
230      <section>
231        <section>
232          <h2>4. Front-end Annotation Interface</h2>
233          <img src="figure/frontend_all.png" alt="Frontend Interface" >
234        </section>
235        
236        <section>
237          <h3>4.1 Worktable</h3>
238          <div class="two-column">
239            <div style="flex: 0.5;">
240              <div style="margin-bottom: 1rem;">
241                <p class="small-text">Upload & Cut Dialogues</p>
242              </div>
243            </div>
244            <!-- 前端界面截图 -->
245            <div style="flex: 1;">
246              <img src="figure/frontend_upload.jpg" alt="Frontend Interface" class="slide-image">
247            </div>
248          </div>
249        </section>
250
251        <section>
252          <h3>4.2 Interactive Annotation</h3>
253          <div class="two-column">
254            <div style="flex: 0.5;">
255              <div>
256                <p class="small-text">Refine annotations: click to switch roles</p>
257              </div>
258            </div>
259            <!-- 前端界面截图 -->
260            <!-- <div style="flex: 1;">
261              <img src="figure/frontend_annotation.png" alt="Frontend Interface" class="slide-image">
262            </div> -->
263            <div style="flex: 1;">
264              <video src="figure/annotation.mp4" controls autoplay muted loop playsinline class="slide-image">
265                Your browser does not support the video tag.
266              </video>
267            </div>
268          </div>
269        </section>
270
271        <!-- <section>
272          <video
273            src="figure/annotation.mp4"
274            controls
275            autoplay
276            loop
277            muted
278            playsinline
279            style="max-width: 100%; height: auto; ">
280          </video>
281        </section> -->
282
283        <section>
284          <h3>4.3 Visualization Features</h3>
285          <div class="two-column">
286            <div style="flex: 1;">
287              <h5>Data visualization design</h5>
288              <ul class="ssmall-text">
289                <li>Large nodes (green) = main name</li>
290                <li>Small nodes (yellow) = aliases</li>
291                <li>Links between large nodes = relations (measured with numbers of conversation)</li>
292                <li>Chapter-specific networks: generating dialogue networks for all character relationships up to the current chapter, reveal
292ing narrative progression.</li>
293              </ul>
294            </div>
295            <div style="flex: 1;">
296              <video src="figure/network.mp4" controls autoplay muted loop playsinline class="slide-image">
297                Your browser does not support the video tag.
298              </video>
299            </div>
300          </div>
301        </section>
302
303        <section>
304          <h3>4.3 Visualization Features</h3>
305          <div class="two-column">
306            <div style="flex: 0.5;">
307              <h5>Manual Disambiguation</h5>
308              <p class="ssmall-text">Correct extraction errors</p>
309            </div>
310            <div style="flex: 1;">
311              <video src="figure/disambiguation.mp4" controls autoplay muted loop playsinline class="slide-image">
312                Your browser does not support the video tag.
313              </video>
314            </div>
315          </div>
316        </section>
317      </section>
318
319      <!-- 5. Case Studies -->
320      <section>
321        <section>
322          <h2>5. Simple Insights from the Data Visualization</h2>
323        </section>
324        
325        <!-- <section>
326          <h3>5.1 Dream of the Red Chamber</h3>
327          <div class="two-column">
328            <div>
329              <h4>Initial Extraction</h4>
330              <p class="small-text">LLM-based character identification</p>
331            </div>
332            <div>
333              <h4>Manual Refinement</h4>
334              <p class="small-text">Human expertise correction</p>
335            </div>
336          </div>
337            <img src="figure/frontend_annotation.png" alt="case" style="width: 70%;">
338        </section> -->
339
340        <section>
341          <!-- <h3>Data Visualization</h3> -->
342          <div class="two-column">
343            <div>
344              <h4>Key Characters</h4>
345              <p class="small-text">Important intermediary characters identified</p>
346            </div>
347            <div>
348              <h4>Interaction Patterns</h4>
349              <p class="small-text">Social dynamics revealed</p>
350            </div>
351          </div>
352          <img src="figure/chapters.png" alt="case" style="width: 70%;">
353          
354        </section>
355        <section>
356          <h4>Network of the first chapter</h4>
357          <img src="figure/first_chapter_net.png" alt="case" style="width: 100%;">
358        </section>
359        <section>
360          <h4>Network of the second chapter</h4>
361          <img src="figure/second_chapter_net.png" alt="case" style="width: 100%;">
362        </section>
363        <section>
364          <h4>Network of the third chapter</h4>
365          <img src="figure/third_chapter_net.png" alt="case" style="width: 100%;">
366        </section>
367        <section>
368          <h4>Network of the forth chapter</h4>
369          <img src="figure/forth_chapter_net.png" alt="case" style="width: 100%;">
370        </section>
371      </section>
372
373      <!-- 6. Conclusion -->
374      <section>
375        <section>
376          <h2>6. Conclusion</h2>
377        </section>
378        
379        <section>
380          <h3>6.1 Summary</h3>
381          <div class="two-column">
382            <div>
383              <p class="small-text">AI + Human</p>
384            </div>
385            <img src="figure/pipeline.png" alt="case" style="width: 70%;">
386          </div>
387        </section>
388
389        <section>
390          <h3>6.2 TODOs</h3>
391          <div class="two-column">
392            <div>
393              <h4>Improving Algorithm Efficiency</h4>
394              <p class="small-text">Improve the speed of data processing, thus generate the result faster.</p>
395            </div>
396            <div>
397              <h4>Incorporating SNA Algorithm</h4>
398              <p class="small-text">Use social network analysis algorithm to build a final network, which can directly serve for the literary analysis.</p>
399            </div>
400          </div>
401        </section>
402      </section>
403
404      <!-- Thank You -->
405      <section>
406        <h2>Thank You!</h2>
407        <p class="small-text" style="font-style: italic;">[email protected]</p>
408        <p class="small-text" style="font-style: italic;">https://yutong-yang.github.io/</p>
409      </section>
410    </div>
411  </div>
412  
412<script src="https://cdn.jsdelivr.net/npm/[email protected]/dist/reveal.js"></script>
412
413  
413<script>
414    Reveal.initialize({
415      hash: true,
416      transition: 'slide'
417    });
418  </script>
418
419</body>
420</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.