PageSourceSearch

https://pcarbo.github.io/objrecls/index.html

html pcarbo.github.io collected 2026-10-03 09:44:42 UTC 12,387 bytes, 297 lines download raw bytes

1<html>
2<title>Learning to recognize objects with little supervision</title>
3<body bgcolor="#FFFFFF">
4
5<center>
6<br>
7<table align="center" border=0 width=460 cellspacing=0
8cellpadding=0>
9<tr><td valign=top>
10
11<a name="title">
12<h2>Learning to recognize objects with little supervision</h2>
13
14<img alt="Pictures of cars" src="cars.gif" width=460 height=134><br>
15
16<p>This is the project webpage that accompanies the journal submission
17with the same name. See <a href="#pub">below</a> for published papers
18relevant to this project. Here you will find the <a
19href="#code">code</a> and <a href="#data">data</a> used in the
20experiments. This webpage is maintained by <a
21href="../index.html">Peter Carbonetto</a>.</p>
22
23<h3>People</h3>
24
25<p>The following people were involved in this project: <a
26href="../index.html">Peter Carbonetto</a>, <a
27href="http://lear.inrialpes.fr/people/dorko">Gyuri Dork&ograve;</a>,
28<a href="http://lear.inrialpes.fr/people/schmid">Cordelia Schmid</a>, <a
29href="http://www.cs.ubc.ca/&#126;nando">Nando de Freitas</a> and <a
30href="http://www.cs.ubc.ca/&#126;kueck">Hendrik K&uuml;ck</a>.</p>
31
32<a name="data">
33<h3>Data</h3>
34
35<p>We used six different databases to evaluate our proposed Bayesian
36model. Five of them were collected and made publicly available by
37other researchers. They are: <a
38href="http://www.robots.ox.ac.uk/~vgg/data3.html">airplanes</a>, <a
39href="http://www.robots.ox.ac.uk/~vgg/data3.html">motorbikes</a>,
40wildcats, <a
41href="http://www.emt.tugraz.at/~pinz/data/GRAZ_01/">bicycles</a> and
42<a href="http://www.emt.tugraz.at/~pinz/data/GRAZ_01/">people</a>. The
43wildcats database is not publicly available since it was created using
44the commercial Corel images database. We created the sixth and final
45data set. It consists of photos of parking lots and cars near the
46INRIA Rh&ocirc;ne-Alpes research centre in Montbonnot, France. The
47INRIA car database is available for download <a
48href="data/cars.tar.gz">here</a>. We now describe how to read in and
49use it.</p>
50
51<p>In the root directory, there are two files <b>trainimages</b> and
52<b>testimages</b>. Each one contains a list of image names, one per
53line. All the images are contained, appropriately enough, in the
54<b>images</b> folder with the file names appended with the <b>.jpg</b>
55extension.</p>
56
57<p>In addition, we have provided manual annotations of the scenes
58which consist of boxes (&quot;windows&quot;) that surround the
59objects. The files describing the scene annotations are contained in
60the <b>objects</b> subdirectory, appended with a <b>.pgm.objects</b>
61suffix.</p> Each line describes a single window and looks like
62this:<center><br><tt>Object: x y width height</tt><br>
63</center></br> where (x,y) is the top-left corner of the box. Note
64that the coordinate system starts at 0, not 1 as in Matlab. Here is a
65function <a href="matlab/loadobjectwindows.m">loadobjectwindows</a>
66for loading the windows from a objects file into Matlab.
67
68<a name="code">
69<h3>Code</h3>
70
71Our approach consists of three steps. First, we use detector to
72extract a sparse set of a priori informative interest regions. Second,
73we train the Bayesian classification model using a Markov Chain Monte
74Carlo algorithm. Third, for object localization, we run inference on a
75conditional random field. The code for the three steps is described
76and made available for download below.
77
78<h4>Interest region detectors</h4>
79
80The three detectors we employed were developed elsewhere. Binaries for
81the Harris-Laplace and Laplacian of Gaussian interest region detectors
82are available for download <a
83href="http://lear.inrialpes.fr/people/dorko/downloads.html">here</a>. You
84can find the Kadir-Brady entropy detector at Timor Kadir's <a
85href="http://www.robots.ox.ac.uk/~timork">website</a>. Once the
86regions are extracted, you still need a way of describing the regions
87in a way that our model will understand. We use the Scale Invariant
88Feature Transform (SIFT) descriptor. With our parameter settings, each
89interest region ends up as a 128-dimension feature vector. The <a
90href="http://lear.inrialpes.fr/people/dorko/downloads.html">same
91package</a> as above can be used to compute the SIFT feature vectors.
92
93<a name="ssmcmc">
94<h4>MCMC algorithm for Bayesian classification</h4>
95
96<p>We have a C implementation of the Markov Chain Monte Carlo (MCMC)
97algorithm for simulating the posterior of the Bayesian kernel machine
98classifier given some training data. It was tested in Linux. Here is
99how to compile and install the code.</p>
100
101<p>First, you need to install a copy of the <a
102href="http://www.gnu.org/software/gsl">GNU Scientific Library</a>
103(GSL). Our code was testted with GSL version 1.6. The libraries should
104be installed in the directory <b>$HOME/gsl/lib</b> and the header
105files in <b>$HOME/gsl/include</b>, and the variable <b>$HOME</b> must
106be entered correctly in the Makefile (see below). In addition, if you
107use the gcc compiler in Linux you have to set the path to include the
108installed libraries with the command<center><tt>setenv
109LD_LIBRARY_PATH $HOME/gsl/lib</tt></center></p>
110
111<p>Next, you're ready to install the <a
112href="c/ssmcmc.tar.gz">ssmcmc</a> package. &quot;ssmcmc&quot; stands
113for &quot;semi-supervised MCMC.&quot; As mentioned, you have to edit
114the file <b>Makefile</b> and make sure that the <b>HOME</b> variable
115points to the right directory. Once you're in the <b>ssmcmc</b>
116directory, type <tt>make</tt>, and after a few seconds you should have
117a program called <b>ssmcmc</b>. Running the program without any input
118arguments gives to the help. We give a brief tutorial explaining how
119to use the program.<p>
120
121<p>In order to train the model on some data, you need a few
122ingredients. First, you need some data in the proper format. We've
123made a sample <a href="data/carhartrain.gz">training set</a> available
124for download. In fact, this particular data set was used for many of
125our experiments. It was produced by extracting Harris-Laplace interest
126regions from the INRIA car data set (an average of 100 regions per
127image) then converting them to feature vectors using SIFT. The format
128of the data is as follows:
129<ul>
130
131<li>The first line gives the number of documents (images).
132
133<li>The second line gives the dimension of the feature vectors.
134
135<li>After that there's a line for each document (image) in the data
136set. Each line has two numbers. The first gives the image caption. It
137can either be 1 (all the points in the document are positive, which
138almost never happens in our data sets), 2 (all the points in the
139document are negative, which happens when there is no instance of the
140object in the image), or 0 (the points are unlabeled, which happens
141when there is an instance of the object in the image). The second
142number says how many points (extracted interest regions) there are in
143the document.
144
145<li>The last part of the data file, and the biggest, is the data
146points themselves. There is one feature vector for each line. The
147first number is the true label --- this is only used for evaluation
148purposes and is not available to the model. The rest of the numbers
149are the entries that make up the feature vector.
150
151</ul></p>
152
153<p>Suppose you have your data set available. Next, you need to specify
154the parameters for the model in a text file. A sample parameters file
155looks like this:<br>
156
157<pre>
158Put comment here.
159ns:      1000
160metric:  fdist2
161kernel:  kgaussian
162lambda:  0.01
163mu:      0.01
164nu:      0.01
165a:       1.0
166b:       50.0
167mua:     0.01
168nua:     0.01
169epsilon: 0.1
170nc1:     30
171nc2:     0
172</pre>
173
174Most of the parameters above are explained in the journal paper
175submission and technical report. <b>ns</b> is the number of samples to
176generate. There is only one possible distance metric and kernel, but
177they must be specified anyway. <b>lambda</b> is the kernel scale
178parameter and <b>epsilon</b> is the stabilization term on the
179covariance prior.</p>
180
181<p>The parameters <b>nc1</b> and <b>nc2</b> specify the minimum number
182of positive labels and the minimum number of negative labels in a
183training image, respectively, so this obviously specifies a
184constrained data association model. In this case, the constraints
185require that at least 30 interest regions in a training image be
186labeled as positive. Alternatively, one can specify data association
187problem using group statistics, in which case the last two lines are
188replaced by something like<br>
189
190<pre>
191m:       0.3
192chi:     400
193</pre></p>
194
195<p>Once the model parameters are specified, you can finally train the
196model with the following command:<br>
197
198<center><tt>
199ssmcmc -t=params -v carhartrain model
200</tt></center><br>
201
202We're assuming here at that the parameters file is called
203<b>params</b>. The result is saved in the file <b>model</b>. Once
204training is complete (it might take a little while), you can use the
205model to predict the labels of interest regions extracted from any
206image, including this <a href="data/carhartest.gz">sample test set</a>,
207with the following command:<br>
208<br>
209
210<center><tt>
211ssmcmc -p=labels -b=100 -v carhartest model
212</tt></center><br>
213
214This specifies a burn-in of 100, so the first hundred samples are
215discarded. The resulting predictions, along with the level of
216confidence in those predictions, is saved in the file
217<b>labels</b>. Each line in the file has two numbers: the probability
218of a positive classification (the interest region belongs to the
219object) and the probability of a negative classification (the interest
220region is associated with the background). These two numbers should
221add up to one. Here is a couple examples from that sample test set.<br>
222
223<br><img alt="Label estimated in a couple test images"
224src="carspredict.gif" width=460 height=174><br><br>
225
226The blue interest regions are more likely to belong to a car (greather
227than 0.5 chance).</p>
228
229<h4>Conditional random field for localization</h4>
230
231<p>We implemented the CRF model for localization in Matlab. The function
232is called <a href="matlab/crflocalize.m">crflocalize</a>. Type <b>help
233crflocalize</b> in the Matlab command line to get instructions on how to 
234use it.
235
236<p>The function crflocalize requires the optimized Matlab
237implementation of the random schedule tree sampler, <a
238href="http://www.cs.ubc.ca/~pcarbo/wiki/bgsfast.tar.gz">bgsfast</a>. Once
239you have downloaded and unpacked the tar ball, follow these steps:
240<ol>
241
242<li>Make sure that you have the GNU Scientific
243Library installed and that LD_LIBRARY_PATH is set correctly (see the
244instructions above). 
245
246<li>Edit the Makefile. You want the variable <b>GSLHOME</b> to point
247to the proper location.
248
249<li>Type <tt>make</tt> in the code directory to compile the MEX files.
250To make sure that it is working, you can run the <b>testbgs</b> Matlab
251script. If you have problems compiling the C code, it may be because
252your MEX options are not set correctly. In particular, make sure you
253are using the g++ copmiler (or another C++ compiler). Refer to the
254Mathworks support website for details. Note that this code has only
255been tested in Matlab version 7.0.1.
256
257<li>Once you've managed to compile the bgsfast MEX program
258successfully, use the <b>addpath</b> function to include the
259<b>bgsfast</b> directory in your Matlab path.
260
261</ol></p>
262
263<p>Here are a couple of localization results for the cars database.<br><br>
264<img alt="Localization of cars" src="carslocalize.gif" width=460 height=105>
265</p>
266
267<h3>Note</h3>
268
269If you have any questions or problems to report about this project, do
270not hesitate to contact the <a href="../index.html">main author</a>.
271
272<a name="pub">
273<h3>Publications</h3>
274
275<p>Hendrik K&uuml;ck and Nando de Freitas. <a
276href="http://www.cs.ubc.ca/&#126;kueck/papers/KueckUAI05.pdf">Learning
277to classify individuals based on group statistics.</a> Conference on
278Uncertainty in Artificial Intelligence, July 2005.</p>
279
280<p>Peter Carbonetto, Gyuri Dork&ograve; and Cordelia Schmid.  <a
281href="../techreport.pdf">Bayesian learning for weakly supervised
282object classification.</a> Technical Report, INRIA Rh&ocirc;ne-Alpes,
283July 2004.</p>
284
285<p>Hendrik K&uuml;ck, Peter Carbonetto and Nando de Freitas.  <a
286href="../semisup.pdf">A Constrained semi-supervised learning approach
287to data association</a>. European Conference on Computer Vision, May
2882004.</p>
289
290<center>
291<hr noshade width="100%" size=1>
292<font color="#333333" size=2>
293This webpage was last updated on August 13, 2005. 
294<a href="../index.html">Home</a>.  </font> </center>
295</td></tr></table>
296</body>
297</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.