PageSourceSearch

https://arjung128.github.io/svm/

html arjung128.github.io collected 2026-10-03 08:39:15 UTC 8,506 bytes, 256 lines download raw bytes

1<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN"
2    "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
3<!-- saved from url=(0069)http://www.cs.cmu.edu/~dchaplot/projects/neural-topological-slam.html -->
4
5<html xmlns="http://www.w3.org/1999/xhtml">
6<!-- ======================================================================= -->
7
8<head>
9  <meta http-equiv="Content-Type" content="text/html; charset=us-ascii" />
10  
10<script type="text/javascript" id="www-widgetapi-script" src="./website_files/www-widgetapi.js" async="">
11</script>
11
12  
12<script src="./website_files/jsapi" type="text/javascript">
13</script>
13
14  
14<script type="text/javascript">
15//<![CDATA[
16  google.load("jquery", "1.3.2");
17  //]]>
18  </script>
18
19  
19<script type="text/javascript" charset="UTF-8" src="./website_files/jquery.min.js">
20</script>
20
21  <style type="text/css">
22/*<![CDATA[*/
23  body {
24    font-family: "Titillium Web","HelveticaNeue-Light", "Helvetica Neue Light", "Helvetica Neue", Helvetica, Arial, "Lucida Grande", sans-serif;
25    font-weight:300;
26    font-size:18px;
27    margin-left: auto;
28    margin-right: auto;
29    width: 1100px;
30  }
31
32  h1 {
33    font-weight:300;
34  }
35
36  .disclaimerbox {
37    background-color: #eee;
38    border: 1px solid #eeeeee;
39    border-radius: 10px ;
40    -moz-border-radius: 10px ;
41    -webkit-border-radius: 10px ;
42    padding: 20px;
43  }
44
45  video.header-vid {
46    height: 140px;
47    border: 1px solid black;
48    border-radius: 10px ;
49    -moz-border-radius: 10px ;
50    -webkit-border-radius: 10px ;
51  }
52
53  img.header-img {
54    height: 140px;
55    border: 1px solid black;
56    border-radius: 10px ;
57    -moz-border-radius: 10px ;
58    -webkit-border-radius: 10px ;
59  }
60
61  img.rounded {
62    border: 1px solid #eeeeee;
63    border-radius: 10px ;
64    -moz-border-radius: 10px ;
65    -webkit-border-radius: 10px ;
66  }
67
68  a:link,a:visited
69  {
70    color: #1367a7;
71    text-decoration: none;
72  }
73  a:hover {
74    color: #208799;
75  }
76
77  td.dl-link {
78    height: 160px;
79    text-align: center;
80    font-size: 22px;
81  }
82
83  .layered-paper-big { /* modified from: http://css-tricks.com/snippets/css/layered-paper/ */
84    box-shadow:
85            0px 0px 1px 1px rgba(0,0,0,0.35), /* The top layer shadow */
86            5px 5px 0 0px #fff, /* The second layer */
87            5px 5px 1px 1px rgba(0,0,0,0.35), /* The second layer shadow */
88            10px 10px 0 0px #fff, /* The third layer */
89            10px 10px 1px 1px rgba(0,0,0,0.35), /* The third layer shadow */
90            15px 15px 0 0px #fff, /* The fourth layer */
91            15px 15px 1px 1px rgba(0,0,0,0.35), /* The fourth layer shadow */
92            20px 20px 0 0px #fff, /* The fifth layer */
93            20px 20px 1px 1px rgba(0,0,0,0.35), /* The fifth layer shadow */
94            25px 25px 0 0px #fff, /* The fifth layer */
95            25px 25px 1px 1px rgba(0,0,0,0.35); /* The fifth layer shadow */
96    margin-left: 10px;
97    margin-right: 45px;
98  }
99
100
101  .layered-paper { /* modified from: http://css-tricks.com/snippets/css/layered-paper/ */
102    box-shadow:
103            0px 0px 1px 1px rgba(0,0,0,0.35), /* The top layer shadow */
104            5px 5px 0 0px #fff, /* The second layer */
105            5px 5px 1px 1px rgba(0,0,0,0.35), /* The second layer shadow */
106            10px 10px 0 0px #fff, /* The third layer */
107            10px 10px 1px 1px rgba(0,0,0,0.35); /* The third layer shadow */
108    margin-top: 5px;
109    margin-left: 10px;
110    margin-right: 30px;
111    margin-bottom: 5px;
112  }
113
114  .vert-cent {
115    position: relative;
116      top: 50%;
117      transform: translateY(-50%);
118  }
119
120  hr
121  {
122    border: 0;
123    height: 1px;
124    background-image: linear-gradient(to right, rgba(0, 0, 0, 0), rgba(0, 0, 0, 0.75), rgba(0, 0, 0, 0));
125  }
126
127  #authors td {
128    padding-bottom:5px;
129    padding-top:30px;
130  }
131  /*]]>*/
132  </style><!-- ======================================================================= -->
133  <!-- Global site tag (gtag.js) - Google Analytics -->
134
135  
135<script async="" src="./website_files/js" type="text/javascript">
136</script>
136
137  
137<script type="text/javascript">
138//<![CDATA[
139  
vendor: 134 bytes, lines 139-143
139window.dataLayer = window.dataLayer || [];
140  function gtag(){dataLayer.push(arguments);}
141  gtag('js', new Date());
142
143  gtag('config', '
143UA-114291442-6
vendor: 6 bytes, lines 143-144
143');
144  
144//]]>
145  </script>
145
146  
146<script type="text/javascript" src="./website_files/hidebib.js">
147</script>
147
148  <link href="./website_files/css" rel="stylesheet" type="text/css" />
149  <meta http-equiv="X-UA-Compatible" content="IE=edge" />
150  <link rel="icon" type="image/png" href="" />
151
152  <title>Estimating Perceptual Uncertainty to Predict Robust Motion Plans</title>
153  <meta name="HandheldFriendly" content="True" />
154  <meta name="viewport" content="width=device-width, initial-scale=1.0" />
155  
155<script src="./website_files/iframe_api" type="text/javascript">
156</script>
156
157</head>
158
159<body>
160  <br />
161
162
163  <center>
164    <span style="font-size:44px;font-weight:bold;">Precise Mobile Manipulation<br/>of Small Everyday Objects</span>
165  </center>
166  <br />
167
168
169  <table align="center" width="800px">
170    <tbody>
171      <tr>
172        <td align="center" width="230px">
173          <center>
174            <span style="font-size:22px">Arjun Gupta</span>
175          </center>
176        </td>
177
178        <td align="center" width="230px">
179          <center>
180            <span style="font-size:22px">Rishik Sathua</span>
181          </center>
182        </td>
183
184        <td align="center" width="230px">
185          <center>
186            <span style="font-size:22px">Saurabh Gupta</span>
187          </center>
188        </td>
189      </tr>
190
191
192      <tr>
193        <td align="center" width="230">
194          <center>
195            <span style="font-size:20px">UIUC</span>
196          </center>
197        </td>
198
199        <td align="center" width="230">
200          <center>
201            <span style="font-size:20px">UIUC</span>
202          </center>
203        </td>
204
205        <td align="center" width="230px">
206          <center>
207            <span style="font-size:20px">UIUC</span>
208          </center>
209        </td>
210      </tr>
211    </tbody>
212  </table>
213
214  <center>
215      <div style="margin-top: 10px; margin-bottom: 10px;">
216          <span style="font-size:25px; font-weight:bold;">RA-L 2026</span>
217      </div>
218  </center>
219
220  <table align="center" width="700px" style="padding-top: 15px">
221    <tbody><tr>
222      <td align="center" width="200px"><center><span style="font-size:24px"><a href="https://arxiv.org/abs/2502.13964">Paper</a></span></center></td>
223    </tr><tr>
224  </tr></tbody></table>
225
226  <table align="center" width="300px" style="padding-top: 25px">
227    <tbody>
228      <tr>
229        <td align="center" width="300px"><iframe width="700" height="400" src="https://www.youtube.com/embed/1BFclCB91G0" frameborder="0" allow="accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe></iframe>
230        </td>
231      </tr>
232    </tbody>
233  </table>
234  <br />
235
236  <br />
237
238
239  <div style="width:800px; margin:0 auto; text-align:justify">
240    Many everyday mobile manipulation tasks require precise interaction with small objects, such as grasping a knob to open a cabinet or pressing a light switch. In this paper, we develop Servoing with Vision Models (SVM), a closed-loop framework that enables a mobile manipulator to tackle such precise tasks involving the manipulation of small objects. SVM uses state-of-the-art vision foundation models to generate 3D targets for visual servoing to enable diverse tasks in novel environments. Naively doing so fails because of occlusion by the end-effector. SVM mitigates this using vision models that out-paint the end-effector, thereby significantly enhancing target localization. We demonstrate that aided by out-painting methods, open-vocabulary object detectors can serve as a drop-in module for SVM to seek semantic targets (e.g. knobs) and point tracking methods can help SVM reliably pursue interaction sites indicated by user clicks. We conduct a large-scale evaluation spanning experiments in 10 novel environments across 6 buildings including 72 different object instances. SVM obtains a 71% zero-shot success rate on manipulating unseen objects in novel environments in the real world, outperforming an open-loop control method by an absolute 42% and an imitation learning baseline trained on 1000+ demonstrations also by an absolute success rate of 50%.
241
242  </div>
243
244  <div style="width:1000px; margin:0 auto; text-align:center">
245     
246  </div>
247
248  <br />
249
250  <hr />
251
252  
252<script xml:space="preserve" language="JavaScript" type="text/javascript">
253  <!--hideallbibs();-->
254  </script>
254
255</body>
256</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.