PageSourceSearch

https://jerrygaolondon.github.io/projects.html

html jerrygaolondon.github.io collected 2026-10-03 09:22:52 UTC 24,275 bytes, 224 lines download raw bytes

1<!DOCTYPE HTML>
2<!--
3	Ex Machina by TEMPLATED
4    templated.co @templatedco
5    Released for free under the Creative Commons Attribution 3.0 license (templated.co/license)
6-->
7<html>
8    <head>
9        <title>JIE GAO</title>
10        <meta http-equiv="content-type" content="text/html; charset=utf-8" />
11        <meta name="description" content="Jie Gao's Home Page" />
12        <meta name="keywords" content="Jie Gao, Jerry Gao, Researcher, NLP, Logically, Sheffield, Southampton" />
13        <link href='https://fonts.googleapis.com/css?family=Roboto+Condensed:700italic,400,300,700' rel='stylesheet' type='text/css'>
14        <!--[if lte IE 8]>
14<script src="js/html5shiv.js"></script>
14<![endif]-->
15        
15<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.11.0/jquery.min.js"></script>
15
16        
16<script src="js/skel.min.js"></script>
16
17        
17<script src="js/skel-panels.min.js"></script>
17
18        
18<script src="js/init.js"></script>
18
19        <noscript>
20            <link rel="stylesheet" href="css/skel-noscript.css" />
21            <link rel="stylesheet" href="css/style.css" />
22            <link rel="stylesheet" href="css/style-desktop.css" />
23        </noscript>
24        <link rel="stylesheet" href="css/font-awesome.css" />
25
26        <!--[if lte IE 8]><link rel="stylesheet" href="css/ie/v8.css" /><![endif]-->
27        <!--[if lte IE 9]><link rel="stylesheet" href="css/ie/v9.css" /><![endif]-->
28		<!-- Global site tag (gtag.js) - Google Analytics -->
vendor: 67 bytes, lines 28-29
28
29		<script async src="https://www.googletagmanager.com/gtag/js?id=
29UA-10404238-8
vendor: 15 bytes, lines 29-30
29"></script>
30		
30<script>
31		  
vendor: 144 bytes, lines 31-35
31window.dataLayer = window.dataLayer || [];
32		  function gtag(){dataLayer.push(arguments);}
33		  gtag('js', new Date());
34
35		  gtag('config', '
35UA-10404238-8
vendor: 7 bytes, lines 35-36
35');
36		
36</script>
36
37
38    </head>
39    <!-- Google tag (gtag.js) -->
vendor: 69 bytes, lines 39-40
39
40    <script async src="https://www.googletagmanager.com/gtag/js?id=
40G-EKD498PHM1
vendor: 17 bytes, lines 40-41
40"></script>
41    
41<script>
42      
vendor: 150 bytes, lines 42-46
42window.dataLayer = window.dataLayer || [];
43      function gtag(){dataLayer.push(arguments);}
44      gtag('js', new Date());
45
46      gtag('config', '
46G-EKD498PHM1
vendor: 9 bytes, lines 46-47
46');
47    
47</script>
47
48	<body class="no-sidebar">
49
50	<!-- Header -->
51		<div id="header">
52			<div class="container">
53					
54				<!-- Logo -->
55					<div id="logo">
56						<h1><a href="#">Jie Gao</a></h1>
57					</div>
58				
59				<!-- Nav -->
60					<nav id="nav">
61						<ul>
62							<li><a href="index.html">Home</a></li>
63                            <li class="active"><a href="projects.html">Projects</a></li>
64							<li><a href="publications.html">Publications</a></li>
65                            <li><a href="resource.html">Resources</a></li>
66                            <li><a href="teaching.html">Teaching</a></li>
67						</ul>
68					</nav>
69
70			</div>
71		</div>
72	<!-- Header -->
73		
74	<!-- Banner
75		<div id="banner">
76			<div class="container">
77			</div>
78		</div> -->
79	<!-- /Banner -->
80
81	<!-- Main -->
82		<div id="page">
83				
84			<!-- Main -->
85			<div id="main" class="container">
86            <div class="row">
87					<div class="12u">
88                    
89                    <section>
90
91
92                    </section>
93                    
94                    <section>
95                        <h2>Past research Projects:</h2><br>
96                        <ul>
97                            <li> <a href="https://www.researchgate.net/project/Early-Rumour-Detection-on-Social-Media-with-weak-supervision-techniques" title="Early Rumor Detection research GATE project page">Social Media Early Rumour Detection and Weak Supervision</a>
98                            </li>
99                            <li>This is a research project that work towards early rumour detection. Our Rumour detection task is to identify pieces of information on social media that is need to be verified.
100                                Manual annotation of large-scale and noisy social media data for rumors is highly labor-intensive, time-consuming and requires special skills and insight to a specific event.
101                                Existing rumor dataset are in relatively small size and suffer high class imbalance. The challenge of data scarcity slows down the progress of machine learning based rumor detection technique.
102                                This project investigate data augmentation techniques in rumor detection for the purpose of exploiting unlabeled social media data to augment limited labeled rumor data.
103                                We are sharing our augmented data to the community to support further development of machine learning based early rumor detection and study of rumor propagation patterns.
104                                In addition, based on the data augmentation technique, we are also working on a neural approach to the task of rumor detection can benefit from the inclusion of handcrafted features and state-of-the-art sentence embeddings.
105                                We present a novel neural network architecture that based on a combination of stacked Long short-term memory (LSTM) network and attention mechanisms that are used to encode tweet social context, and a SoA character-based biLM language model to encode short-text tweet.
106                                The main purpose of the research is to research and evaluate SoA multimodal feature representation learning through DNN and attempt to understand how attention mechanism helps to handle noisy context in social media.
107                                We have conducted ablation study to understand relative contribution of each component of our proposed model.
108                                Our neural approach achieves state-of-the-art performance on a collection of public dataset with strict evaluation and further improved marginally with a large augmented rumor dataset.
109                                We are now working on publishing our findings along with source code and dataset used in our experiment.
110                            </li>
111                        </ul>
112                        <ul>
113                            <li>
114                                <span class="fa featured"><a href="https://www.europeana.eu/portal/en" title="Europeana website"> <img class="image-project-logo" alt="Europeana logo" src="images/Europeana-logo.jpg"  WIDTH="220" HEIGHT="90" />
114 </a></span>
115                            </li>
116                            <li> <a href="https://pro.europeana.eu/project/europeana-dsi-3">Europeana DSI-3</a>
117                            </li>
118                            <li>Europeana DSI-3 is a continuation of the previous Europeana DSI projects (Europeana DSI and Europeana DSI-2). The DSI-3 project operates the Europeana core service platform from mid-2017 to mid-2018. I'm the co-PI of the project. The goal of this project is the development of state-of-the-art (SoA) information retrieval (IR) methodology for the purpose of improving the availability and accessibility for European culture heritage. For current <a href="https://www.europeana.eu/portal/en" title="Europeana Collections Portal">Europeana Collections</a> platform, we are helping to improve indexing,search and infrastructure management for more than <i>50 million</i> digital objects (e.g., artworks, books, videos, sounds) from more than <i>3500</i> museums, galleries, libraries and archives across Europe. In addition to manage current digital objects search platform, the collaboration project mainly focus on digitising <i>18 million</i> historical newspapers from past 200 years which covers <i>10 million</i> full text newspapers issue pages. In this project, we are researching and developing on SoA IR methods and new architecture to address the challenges of handling OCR generated text and NLP/NLU for historical text, meanwhile ensuring cross-language search. The work-in-progress can be tracked via <a href="https://pro.europeana.eu/project/europeana-dsi-3" title="Europeana DSI-3 periodic report">periodic report</a> and our <a href="https://github.com/europeana/search" title="Europeana search infrastructure github repository">github repository</a>.
119                                The good news is that <a href="https://www.europeana.eu/portal/en/collections/newspapers" title="Europeana Newspapers Portal">Europeana Newspapers</a> has been published since 5th December, 2018. You can find the latest <a href="https://europeana.atlassian.net/wiki/spaces/RD/pages/341082156/Infrastructure" title="search infrastructure documentation">documentation</a> about the search infrastructure.
120                            </li>
121                        </ul>
122                        <ul>
123                            <li>
124                                <span class="fa featured"><a href="http://setamobility.eu/" title="SETA project website"> <img class="image-project-logo" alt="SETA logo" src="images/SETA-logo.png"  WIDTH="220" HEIGHT="90" /> </a></span>
125                            </li>
126                            <li> <a href="http://setamobility.eu/">SETA – ubiquitous data and service ecosystem for better metropolitan mobility</a>
127                            </li>
128                            <li>The objective of the SETA project is to provide effective solutions for intelligent and sustainable mobility i.e. the smarter, greener and more efficient movement of people and goods. SETA will provide a radical change from transport as a series of separate modal journeys to an integrated, reactive, intelligent, mobility system. It will provide always-on, pervasive services to citizens and business, as well as decision makers to support safe, sustainable, effective, efficient and resilient mobility. The project lasts 3 years, with €5.5M of funding from EU Horizon 2020 of which €1.2M is for Sheffield. Professor Ciravegna is the project director (2016-2019). I'm one of researchers & developers for the mobility tracking app working with Professor Ciravegna. I'm also the lead developer and manager of SETA mobility data collection infrastructure. Please find our deliverables via <a href="http://setamobility.eu/" title="SETA project website">SETA website</a> and the tracking app is available and will be updated in regular basis via <a href="https://play.google.com/store/apps/details?id=oak.shef.ac.uk.abstractcityactivity">Google Play</a>.
129                            </li>
130                        </ul>
131                        <ul>
132                            <li>
133                                <span class="fa featured"><a href="https://www.nhs.uk/oneyou/active10/home" title="Active10 project website"> <img class="image-project-logo" alt="Active10 logo" src="images/active10-logo.JPG"  WIDTH="220" HEIGHT="90" /> </a></span>
134                            </li>
135                            <li> <a href="https://www.nhs.uk/oneyou/active10/home#h4iMbzX9siYHiA5L.97" title="One you Active10 home page">One You Active 10 Walk Tracker</a>
136                            </li>
137                            <li>This project aims to address adults inactivity by creating a free app to motivate and measure how much brisk walking you are doing throughout the day and highlights how many continuous chunks of 10 minutes – known as Active 10s you achieve.
138                                The project is funded by Public Health England (PHE). I'm one of the major developers working with Prof. Fabio Ciravegna on Android tracking app, leading test and manage regular release and monitoring of PRODUCTION (through <a href="https://play.google.com/apps/publish/" title="Google Play Console link">Google Play Console</a> and <a href="https://get.fabric.io/" title="Fabric - App Development Platform for teams">Fabric</a>). I am the lead developer and manager of Amazon Web Services (AWS) cloud (<a href="aws.amazon.com/ec2‎">EC2</a> and <a href="https://aws.amazon.com/rds/">RDS</a>) based data collection infrastructure (RESTful APIs + Node.js cluster + PM2 + custom python/shell ETL & monitoring tools).  Currently, big data collections are easy to do, but processing and analysing at scale becomes increasingly challenge. The platform handles the large amount of requests and support further data analysis. Data safety and privacy are also paramount. I have also led the work of data security, protection and management with respect to data confidentiality, integrity, availability and risk management. This work has been published via ISCRAM. I am also contributing to the development of the tracking app for Android, leading testing and the production and manage regular release of new versions into production. The Active 10 app has had over 600,000 downloads over 10 months.
139                                See also our <a href="https://www.sheffield.ac.uk/dcs/news/active-10-app-developed-researchers-department-public-health-england-featured-bbc1">department news</a> and <a href="https://www.bbc.co.uk/sport/get-inspired/41024987" title="BBC coverage">BBC coverage</a> for details.
140                            </li>
141                        </ul>
142                        <ul>
143                            <li>
144                                <span class="fa featured"><a href="https://www.movemoresheffield.com/app" title="MoveMore App project website"> <img class="image-project-logo" alt="MoveMore app logo" src="images/move_more_logo.png"  WIDTH="220" HEIGHT="90" /> </a></span>
145                            </li>
146                            <li> <a href="https://www.movemoresheffield.com/app">Move More App</a>
147                            </li>
148                            <li>This project is to help Sheffield become the most active city in the UK by 2020. The app aims to stir some healthy competition. The app is designed to collect and reward even the smallest burst of exercise by counting our Move More Minutes of activity. This makes it easy for anyone to get involved in the challenge, regardless of fitness level. I'm one of two major developers on Android tracking app and university private cloud based data collection infrastructure.
149                            </li>
150                        </ul>
151                        <ul>
152                            <li>
153                                <span class="fa featured"><a href="http://staffwww.dcs.shef.ac.uk/people/F.Ciravegna/Speak-PC/" title="SPEEAK-PC project website"> <img class="image-project-logo" alt="SPEEAK-PC logo" src="images/speeak-pc-log.JPG" /> </a></span>
154                            </li>
155                            <li> <a href="http://staffwww.dcs.shef.ac.uk/people/F.Ciravegna/Speak-PC/" title="SPEEAK-PC project information page">SPEEAK-PC – Sustained Process Excellence through Embedding of Analytics and Knowledge Management into Process Chain</a>
156                            </li>
157                            <li>This is a Innovate UK funded project (project no: 101947). This project directly addresses the challenge for organisations to realise actionable knowledge from an ever increasing flood of potentially valuable data. Currently, the need for skilled ‘data scientists’ is a major bottleneck in this regard. The project applied existing techniques and developed new technologies to create an ICT tool set which alleviates this bottleneck through provision of a collaborative platform with tools for data integration and analytics deployment which are accessible to non-ICT specialists. The project is funded by Innovate UK, start date 1.10.2014 and has a length of 18 months.
158                            I was leading & coordinating the research for work package (WP) 3 (“Data Representation, Mining & Analysis”) and responsible for the delivery for two work packages (WP3 and WP4 – “terminology driven text mining and knowledge discovery application”) in SPEEAK-PC collaboration project with TATA Steel funded by Innovative UK. The deliveries have been successfully released to TATA Steel R&D team for further evaluation and commercialisation. The project was rated as second higher.
159							Please see one of our deliverables -  SPEEAK-PC Terminology Recognition, accessible via the <a href="https://github.com/jerrygaoLondon/SPTR" title="SPTR project github repository"> [link]</a>.
160                            </li>
161                        </ul>
162                        <ul>
163                            <li>
164                                <span class="fa featured"><a href="http://www.wesenseit.com/" title="WeSenseIt project website"> <img class="image-project-logo" alt="wesenseit logo" src="images/wesenseit-logo-small.jpg" WIDTH="220" HEIGHT="90" /> </a></span>
165                            </li>
166                            <li> <a href="http://www.wesenseit.com/">WeSenseIt – Citizen Water Observatories</a>
167                            </li>
168                            <li>WeSenseIt (www.wesenseit.eu) is a multi-site, multi-disciplinary project involving researchers and industrial partners in web technologies, environment, sensing, as well as social media monitoring. Working together with the EU and our sister projects this project aims to develop a new concept of citizen observatories of water creating a two communication channel between authorities and citizens in cooperating to monitor rivers, covering water quality to flooding.
169                            </li>
170                        </ul>
171                        <ul>
172
173                            <li> <a href="http://gow.epsrc.ac.uk/NGBOViewGrant.aspx?GrantRef=EP/J019488/1">Lodie,Web Scale Information Extraction via Linked Open Data</a>
174                            </li>
175                            <li>The linked open data information-extraction (LODIE) project, funded by <a href="https://www.epsrc.ac.uk/" title="EPSRC website">EPSRC</a>, focuses on the study of IE models and algorithms able to perform efficient user-centered web-scale learning by exploiting linked open data (LOD).
176                                I joined the LODIE team in final exploitation stage.
177                            </li>
178                        </ul>
179                        <ul>
180                            <li>
181                                <span class="fa featured"><a href="https://www.justgiving.com" title="JustGiving project website"> <img class="image-project-logo" alt="JustGiving logo" src="images/justgiving-logo_0.png" WIDTH="220" HEIGHT="90" /> </a></span>
182                            </li>
183                            <li> <a href="https://www.justgiving.com">JustGiving</a>
184                            </li>
185                            <li> This is an industrial project collaborated with JustGiving data science team, aiming to apply artificial intelligence algorithms to knowledge mining and user/cause recommendation. The projects is consist of two phrases starting from Oct 2014 to Feb 2015. The main objective of the project ('CauseCat' and 'PeopleCat') and is creating user profiles for JustGiving users, adopting a feature space which is compatible with JustGiving data representation model. We start from identifying relevant data sources which are potentially useful for the categorization of JustGiving users, understand how reliable they are and their (charitable) causes. We analysed sample data obtained from <a href="http://www.experian.co.uk">Experian</a>, as well as data obtainable from various social networks. The output of the exploration phrase is an internal technical report on the analysis of such data sources with an estimation of usefulness to characterise JustGiving users' profiles. In the following phrases, we then focused on identifying the actions needed to align the user profile feature space with JustGiving data representation, and understand if any adjustment is needed on the existing JustGiving data representation model. The output of this phrase is the implementation plan for a global JustGiving data representation (to represent users, charities, causes in the same feature space). In the final phrases, we implemented a prototype for further evaluation in JustGiving PROD. The outcome of the project was finally evaluated in JustGiving PROD to drive donations and engagements via e.g., "You might be interested in" feed card and email subscription. Additional geographical element is added to avoid suggesting local charities to non-local users. The experiment was targeted up to 100k users. The initial A/B testing result shows that our model's output is a <b>statistically significant predictor of what users are interested in</b>. Major techniques in this project cover language modelling, information extraction and clustering, probablistic modelling and supervised learning.
186                            </li>
187                        </ul>
188                        <ul>
189                            <li>
190                                <span class="fa featured">
191                                    <a href="https://www.westminster.ac.uk/business/business-improvement/case-studies/applying-web-intelligence-to-improve-corporate-websites" title="KTP project case study website">
192                                        <img class="image-project-logo" alt="KTP case study logo" src="images/ktp_project.jpg" WIDTH="220" HEIGHT="90" /> </a></span>
193                            </li>
194                            <li> <a href="https://www.westminster.ac.uk/business/business-improvement/case-studies/applying-web-intelligence-to-improve-corporate-websites">KTP Project</a>
195                            </li>
196                            <li> The R&D project (KTP project 007326) is funded by the Innovative UK (formerly UK Technology Strategy Board).
197							I was a leading team member in the project to undertake the Knowledge Transfer Partnership (KTP) objectives as defined in the KTP proposal covering the development of new semantic systems and services on ActiveStandards platform, in order to assist businesses and organizations to discover, understand and act on the actionable knowledge within their information repositories.
198                                This is an industry collaboration project with Magus Research Ltd (now <a href="https://www.crownpeak.com/">Crownpeak</a>) and <a href="https://www.westminster.ac.uk/research/groups-and-centres/cognitive-computing-research-group/people">Cognitive Computing Research Group</a> in University of Westminster.
199
200							Main technologies in this project covers knowledge extraction (more specifically cross-domain NER and relation extraction with <a href="https://gate.ac.uk/" title="GATE website">GATE</a>), knowledge management and retrieval (RDF, RDFS, OWL) with OWLIM-SE (now rebranded as <a href="http://graphdb.ontotext.com/" title="GraphDB website">GraphDB</a>), entity linking and knowlege base enrichment (via linked open data) and Enterprise Information Integration (ontology normalisation & alignment), Web compliance (SEO and fact checking).  For ontology population technique, a upper level ontology is adopted as basis and a starting point for the knowledge modelling of various vertical domains. For the semantic infrastructure, our solution extends <a href="https://ontotext.com/knowledgehub/publications/17141-2/">KIM platform</a> through the collabration with Ontotext R&D team.
201
202								The partnership project has been graded as second higher level which will be shortlisted as one of KTP case-studies among 250 national funded projects. Please see more details in the <a href="https://www.westminster.ac.uk/business/business-improvement/case-studies/applying-web-intelligence-to-improve-corporate-websites" title="KTP project case study page">case study</a> page.
203                            </li>
204                        </ul>
205
206                    </section>
207
208                        </div>
209					</div>
210                    </div>
211			</div>
212			<!-- Main -->
213
214		</div>
215	<!-- /Main -->
216
217
218	<!-- Copyright -->
219    <div id="copyright" class="container">&copy; 2025 Jie Gao, provided under a Creative Commons Attribution license <a href="http://creativecommons.org/licenses/by/3.0/"><img src="images/cc-by.png" height="10" alt="Creative Commons Attribution license"/></a> <br />
220        Template design by <a href="http://templated.co">templated.co</a> </div>
221
222
223	</body>
224</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.