PageSourceSearch

https://hardikmeisheri.github.io/blog/neurips-2019.html

html hardikmeisheri.github.io collected 2026-10-03 09:16:47 UTC 11,958 bytes, 180 lines download raw bytes

1<!DOCTYPE html>
2<html lang="en" data-theme="light">
3<head>
4  <meta charset="UTF-8">
5  <meta name="viewport" content="width=device-width, initial-scale=1.0">
6  <title>NeurIPS 2019 Notes — Hardik Meisheri</title>
7  <meta name="description" content="Conference notes from NeurIPS 2019 — RL talks, NLP highlights, and an informal discussion with Sutton, Silver, and Littman at the RL social.">
8  <link rel="preconnect" href="https://fonts.googleapis.com">
9  <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
10  <link href="https://fonts.googleapis.com/css2?family=Inter:ital,opsz,wght@0,14..32,300;0,14..32,400;0,14..32,500;0,14..32,600;0,14..32,700;1,14..32,400&display=swap" rel="stylesheet">
11  <link rel="stylesheet" href="../assets/css/style.css">
12  <link rel="stylesheet" href="../assets/css/notebook.css">
13</head>
14<body>
15
16<header class="site-header">
17  <nav class="nav container">
18    <a href="../" class="nav-name">Hardik Meisheri</a>
19    <ul class="nav-links" id="nav-links">
20      <li><a href="../#about"        class="nav-link">About</a></li>
21      <li><a href="../#experience"   class="nav-link">Experience</a></li>
22      <li><a href="../#publications" class="nav-link">Publications</a></li>
23      <li><a href="../#contact"      class="nav-link">Contact</a></li>
24      <li><a href="./"               class="nav-link" style="color:var(--accent)">Blog</a></li>
25    </ul>
26    <div class="nav-actions">
27      <button class="theme-toggle" id="theme-toggle" aria-label="Toggle dark mode">
28        <svg class="icon-sun" xmlns="http://www.w3.org/2000/svg" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><circle cx="12" cy="12" r="4"/><path d="M12 2v2M12 20v2M4.93 4.93l1.41 1.41M17.66 17.66l1.41 1.41M2 12h2M20 12h2M6.34 17.66l-1.41 1.41M19.07 4.93l-1.41 1.41"/></svg>
29        <svg class="icon-moon" xmlns="http://www.w3.org/2000/svg" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="M12 3a6 6 0 0 0 9 9 9 9 0 1 1-9-9Z"/></svg>
30      </button>
31      <button class="nav-hamburger" id="nav-hamburger" aria-label="Toggle menu">
32        <span></span><span></span><span></span>
33      </button>
34    </div>
35  </nav>
36</header>
37
38<main>
39<div class="blog-post container">
40
41  <a href="./" class="post-back">
42    <svg xmlns="http://www.w3.org/2000/svg" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="m15 18-6-6 6-6"/></svg>
43    All posts
44  </a>
45
46  <header class="post-header">
47    <h1 class="post-title">NeurIPS 2019 — Conference Notes</h1>
48    <div class="post-meta">
49      <time>December 14, 2019</time>
50      <span class="blog-tag">Conference</span>
51      <span class="blog-tag">RL</span>
52      <span class="blog-tag">NLP</span>
53    </div>
54  </header>
55
56  <div class="post-body">
57
58    <p>I presented my work on Pommerman at the Deep Reinforcement Learning Workshop at NeurIPS this year: <em>"Accelerating training in Pommerman with Imitation and Reinforcement Learning"</em> (with Omkar Shelke, Richa Verma and Harshad Khadilkar). The main highlight for me was the RL social — a new addition by NeurIPS for informal discussions with prominent people in the field. I had a detailed discussion with Richard Sutton, David Silver, Martha White, and Michael Littman over topics ranging from causality in RL to moving away from the MDP framework entirely. The common notion that resonated with all of them: have a big picture in mind before delving into a very specific problem. If building AGI is the ultimate goal, place your work in that context.</p>
59
60    <p>
60In RL, people have been looking at sample-efficient RL, batch RL, meta-learning, and ablation studies of existing algorithms. What follows is a summary of talks and papers I found interesting. Recordings are available at <a href="https://slideslive.com/neurips/" target="_blank" rel="noopener">slideslive.com/neurips/</a>.</p>
61
62    <hr>
63
64    <h2>Tutorials</h2>
65
66    <h3>Imitation Learning and its Application to Natural Language Generation</h3>
67    <p><em>Kyunghyun Cho, Hal Daume III</em></p>
68    <p>Focused on using imitation and reinforcement learning in NMT, dialogue, and story generation. The main challenge with beam search is lack of diversity — typically tackled by adding noise. RL with stochastic policies can help significantly here, especially for natural dialogue generation.</p>
69
70    <h3>Efficient Processing of Deep Neural Networks: from Algorithm to Hardware</h3>
71    <p><em>Vivienne Sze</em></p>
72    <p>Insights into designing efficient hardware for DL under constraints of speed, latency, energy, and cost. Discussion of CPUs, GPUs, FPGAs, and task-specific architectures, mostly focused on vision tasks.</p>
73
74    <h3>Reinforcement Learning: Past, Present and Future Perspectives</h3>
75    <p><em>Katja Hofmann</em></p>
76    <p>An elaborate session from basic MDP formulation to multi-agent RL, with a focus on generalization and policy evaluation. The Minecraft case study was particularly interesting for long-term reward and exploration challenges.</p>
77
78    <hr>
79
80    <h2>Invited Talks &amp; Keynotes</h2>
81
82    <h3>Celeste Kidd: How to Know</h3>
83    <p>Kidd's work centers on how humans (and babies) form beliefs — what they look at, where they focu
83s, what predictions they make. Key takeaways:</p>
84    <ul>
85      <li>Humans continuously form beliefs — it's a probabilistic, ongoing process, not a one-shot decision.</li>
86      <li>Certainty diminishes interest: agents learn most efficiently in the intermediate zone between fully known and fully unknown.</li>
87      <li>Certainty is driven by feedback — without it, our models can be wildly off. Less feedback may encourage overconfidence.</li>
88      <li>Humans form beliefs quickly, meaning the algorithms pushing content online have profound impacts on what we believe.</li>
89    </ul>
90    <p>Her discussion of how small, confirmatory feedback loops solidify wrong assumptions resonated deeply. Standing ovation from the crowd.</p>
91
92    <h3>Yoshua Bengio: From System 1 Deep Learning to System 2 Deep Learning</h3>
93    <p>Inspired by Kahneman's "Thinking Fast and Slow":</p>
94    <ul>
95      <li><strong>System 1</strong>: Intuitive, fast, unconscious, habitual — DL is good at this.</li>
96      <li><strong>System 2</strong>: Slow, logical, sequential, conscious, algorithmic — DL is not equipped for this.</li>
97    </ul>
98    <p>What's missing in DL: out-of-distribution generalization, high-level cognition (causality), and world models that enable knowledge-seeking. The talk proposed moving toward sparse factor graphs and a "consciousness prior."</p>
99
100    <h3>Michael Littman: Assessing the Robustness of Deep RL Algorithms</h3>
101    <p>Built saliency maps by masking portions of the state space (DQN on Atari). Conclusion: DQN does not learn the state space as we see it. Although it has strategies to win games, it doesn't understand them. Evaluation metrics discussed: Value Estimation Error and total accumulated reward.</p>
102
103    <hr>
104
105    <h2>Selected Papers</h2>
106
107    <h3><a href="https://arxiv.org/abs/1907.04595" target="_blank" rel="noopener">Towards Explaining the Regularization Effect of Initial Large Learning Rate</a></h3>
108    <p>Proves that lower learning rates learn myopic/local structure, while larger rates learn macro structure. Demonstrated on CIFAR with a superimposed image patch — high-LR networks learn the macro (generalizable) pattern.</p>
109
110    <h3><a href="https://arxiv.org/abs/1902.04742" target="_blank" rel="noopener">Uniform Convergence May Be Unable to Explain Generalization in Deep Learning</a> (Outstanding New Directions Paper)</h3>
111    <p>Challenges the standard explanation for why over-parameterized networks generalize. Shows that generalization error also depends on training set size in ways uniform convergence cannot capture.</p>
112
113    <h3><a href="https://arxiv.org/abs/1905.11979" target="_blank" rel="noopener">Causal Confusion in Imitation Learning</a></h3>
114    <p>Behavioral cloning fails catastrophically on distributional shift. More information doesn't always help: a policy trained with brake-light information learns to brake whenever the brake light is on — even when it should accelerate. Targeted intervention by expert queries helps resolve this.</p>
115
116    <h3><a href="https://arxiv.org/abs/1902.05546" target="_blank" rel="noopener">Learning to Control Self-Assembling Morphologies</a></h3>
117    <p>Lego-style modules where each takes inputs + messages and produces outputs + messages. Shared policies across modules lead to robust behavior. The <a href="https://youtu.be/cg-RdkPtRiQ" target="_blank" rel="noopener">video</a> is remarkable — the agent completes its task even after losing modules.</p>
118
119    <h3><a href="https://arxiv.org/abs/1906.04358" target="_blank" rel="noopener">Weight Agnostic Neural Networks</a></h3>
120    <p>How much can network structure alone achieve without training? Using random weights with fixed structure, they show that architecture alone yields 82% on MNIST — raising deep questions about the role of inductive biases.</p>
121
122    <h3><a href="https://arxiv.org/abs/1910.12807" target="_blank" rel="noopener">Better Exploration with Optimistic Actor Critic</a></h3>
123    <p>Policy gradient is too greedy, leading to conservative policies. Providing an upper bound (optimistic estimate) to critic values rather than a lower bound leads to better exploration. Improvement over SAC.</p>
124
125    <h3><a href="https://arxiv.org/abs/1912.02503" target="_blank" rel="noopener">Hindsight Credit Assignment</a></h3>
126    <p>Credit assignment is typically done over temporal scale with noisy proxies. This work explicitly learns relevant credit using posterior probabilities — analogous to figuring out why you got wet hours after the fact, rather than just associating it with the most recent actions.</p>
127
128    <hr>
129
130    <h2>Selected Posters</h2>
131
132    <h4>Reinforcement Learning</h4>
133    <ul>
134      <li>Doubly-Robust Lasso Bandit — Kim &amp; Paik, Seoul National University</li>
135      <li>Multi-agent Common Knowledge Reinforcement Learning — Schroeder de Witt et al.</li>
136      <li>Measuring the Reliability of RL Algorithms — Chan et al.</li>
137      <li>Learning Efficient Representations for Intrinsic Motivation — Zhao, Tiomkin, Abbeel</li>
138      <li>Benchmarking Safe Exploration in DRL — Ray, Achiam, Amodei</li>
139    </ul>
140
141    <h4>NLP</h4>
142    <ul>
143      <li>Text-Based Interactive Recommendation via Constraint-Augmented RL — Zhang et al.</li>
144      <li>Can Unconditional Language Models Recover Arbitrary Sentences? — Subramani, Bowman, Cho</li>
145    </ul>
146
147    <h4>Misc</h4>
148    <ul>
149      <li>Adversarial Examples Are Not Bugs, They Are Features — Ilyas et al., Madry group (MIT)</li>
150      <li>Putting an End to End-to-End: Gradient Isolated Learning — Löwe et al.</li>
151    </ul>
152
153  </div>
154</div>
155</main>
156
157<footer class="site-footer">
158  <div class="container">
159    <p>&copy; <span id="footer-year">2026</span> Hardik Meisheri</p>
160  </div>
161</footer>
162
163<script>
164  const THEME_KEY = 'hm-theme';
165  const saved = localStorage.getItem(THEME_KEY);
166  const preferred = window.matchMedia('(prefers-color-scheme: dark)').matches ? 'dark' : 'light';
167  document.documentElement.setAttribute('data-theme', saved || preferred);
168  document.getElementById('theme-toggle').addEventListener('click', () => {
169    const cur = document.documentElement.getAttribute('data-theme');
170    const next = cur === 'dark' ? 'light' : 'dark';
171    document.documentElement.setAttribute('data-theme', next);
172    localStorage.setItem(THEME_KEY, next);
173  });
174  const hamburger = document.getElementById('nav-hamburger');
175  const navLinks  = document.getElementById('nav-links');
176  hamburger.addEventListener('click', () => navLinks.classList.toggle('open'));
177  document.getElementById('footer-year').textContent = new Date().getFullYear();
178</script>
178
179</body>
180</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.