Pierre Sermanet

Pierre Sermanet

Co-founder & Chief Scientist at UMA. Formerly Research Scientist at Google DeepMind and a founding member of Google Brain Robotics. PhD with Yann LeCun at New York University.

Selected highlights
2026
Co-founded UMA
2025
>100k citations across machine learning, computer vision and robotics. Laws of robotics generated from images and sci-fi. CVPR Prize for 10-year significant impact on computer vision by the GoogLeNet paper.
2023
Open-sourced a large-scale VQA benchmark for robotics and developed embodied long-horizon reasoning.
2020
First latent VLA for manipulation: end-to-end vision-language-action models for robotic manipulation, learned from unstructured play videos and very few labels.
2017
Fully self-supervised latent imitation from video — a robot policy learned without any labels.
2016
Founding member of the Google Brain Robotics team. Unsupervised rewards from vision.
2014
Winner of the Dogs vs Cats and 2014 ImageNet competitions. Joined Google Brain.
2013
PhD graduation with Yann LeCun at New York University. Open-sourced one of the first deep vision models (OverFeat). Winner of the 2013 ImageNet competition.
2009
Designed and released an early open-source deep learning framework (EBLearn).
2006
Proposed a fast & slow thinking architecture for robotics (system 1 / system 2).
2005
Deep learning for robotics in DARPA's LAGR program.
2004
Eurobot competition: wrote the vision and control stack of our custom robot.
Filter 33 projects
UMA — Chapter 2 Action UMA — Chapter 2 Planning in latent space, so a robot can simulate trajectories before it moves. 2026 3 Action UMA — Chapter 1 A full humanoid and the embodied AI to drive it, built in nine months by a small team in Paris. 2026 SciFi-Benchmark8 Safety SciFi-Benchmark Robot behaviour measured against science fiction, and constitutions derived from it. 2025 · 2025 ASIMOV Constitutions11 Safety ASIMOV Constitutions A benchmark for semantic safety, with robot constitutions generated from images. 2025 · 2025 9 Reasoning RoboVQA 800k video question-answer pairs for long-horizon robot and human tasks. 2023 · ICRA 2024 · CoRL and NeurIPS 2023 workshops Open X-Embodiment6 Action Open X-Embodiment Robot learning datasets pooled across many embodiments, and the RT-X models trained on them. 2023 · CoRL 2023 RT-29 Action RT-2 Web knowledge carried into closed-loop robot control by co-fine-tuning a vision-language model. 2023 · CoRL 2023 PaLM-E5 Action PaLM-E A language model with continuous sensor observations injected into its embedding space. 2023 · ICML 2023 Inner Monologue7 Reasoning Inner Monologue Language-model planning that replans as grounded feedback arrives from the world. 2022 · CoRL 2022 SayCan7 Reasoning SayCan The language model proposes; learned affordances decide what the robot can actually do. 2022 · CoRL 2022 9 Action Grounding Language in Play Robots commanded in real time by typed natural language, learned from teleoperated play. 2020 · RSS 2021 10 Action Latent Plans from Play Multi-task manipulation self-supervised from cheap unlabelled play data, with no RL. 2019 · CoRL 2019 4 Latent Self-Supervision Actionable Representations Continuous control learned from raw pixels via multi-frame time-contrastive embeddings. 2018 · IROS 2018 10 Latent Self-Supervision Time-Contrastive Networks Self-supervised representations from unlabelled video, rich enough to drive robot imitation. 2017 · ICRA 2018 7 Latent Self-Supervision Perceptual Rewards Reward functions learned without supervision from a handful of human demonstrations. 2017 · RSS 2017 6 Vision Visual Attention A foveated attention RNN whose tracking behaviour emerges from still-image training. 2015 · ICLR 2015 (workshop) Inception / GoogLeNet5 Vision Inception / GoogLeNet The deep architecture that took first place in ImageNet 2014 classification and detection. 2014 · CVPR 2015 · CVPR 2025 ten-year impact prize Dogs vs. Cats Vision Dogs vs. Cats First place in the Kaggle dog-versus-cat image classification challenge. 2014 · 2014 OverFeat10 Vision OverFeat One of the first open deep vision models, and winner of ImageNet 2013 localization. 2013 · ICLR 2014 5 Vision Pedestrian Detection State-of-the-art pedestrian detection using deep ConvNets in EBLearn. 2013 · CVPR 2013 House Numbers6 Vision House Numbers State-of-the-art house-number classification using deep ConvNets. 2012 · ICPR 2012 Traffic Sign Recognition7 Vision Traffic Sign Recognition ConvNets with skip connections, combining low- and high-level features for sign recognition. 2011 · IJCNN 2011 Unsupervised Hierarchies6 Latent Self-Supervision Unsupervised Hierarchies Unsupervised learning of sparse convolutional feature hierarchies that improved a supervised… 2010 · NIPS 2010 EBLearn10 Vision EBLearn An open C++ deep learning framework built on energy-based models. 2009 · ICTAI 2009 2 Robotics NYU Robotics Class Teaching assistant for the NYU robotics course, building the student robot platforms. 2009 · 2009 Robotics LAGR The DARPA programme where ConvNets drove long-range off-road navigation, 2004-2008. 2005 · DARPA, 2004–2008 20 Robotics Long-Range Vision Self-supervised deep vision for long-range autonomous off-road driving. 2009 · JFR 2009 15 Robotics Fast & Slow Navigation Navigation that decouples fast short-range reaction from slow long-range planning. 2009 · JFR 2009 9 Robotics Maneuver Dictionaries Vehicle dynamics learned by recording driven trajectories instead of modelling them. 2008 · ISR 2008 12 Robotics Mapping under Uncertainty A hyperbolic-polar map built to absorb the imprecision of long-range vision. 2008 · IROS 2008 Deep Belief Net Vision11 Robotics Deep Belief Net Vision Self-supervised long-range visual navigation with deep ConvNets. 2008 · IROS 2008 Online Learning7 Robotics Online Learning Long-range vision adapted online, self-supervised by short-range stereo. 2007 · RSS 2007 11 Robotics EUROBOT 2004 Vision-based robot behaviours for a robot-rugby competition. 2004 · 2004

Action · Robotics — 2026

UMA — Chapter 2: Embodied Latent World Models

UMA

Latent world models for humanoids and other robots: planning in latent space, so a robot can simulate trajectories before it moves.

Selected among 2% of startups — one of ten teams — for SPRIND's Next Frontier AI Challenge, a €125M programme from the German Federal Agency for Disruptive Innovation backing teams building new AI paradigms. The grant funds this work.

We're hiring. Get in touch via uma.bot.

Action · Robotics — 2026

UMA — Chapter 1: A full humanoid and embodied AI in nine months

UMA

We built a full humanoid and the embodied AI to drive it in just 9 months, with locomotion and precise manipulation for real-world tasks.

The first prototype, Version 0, integrates AI, software and hardware, all developed from the ground up by a small team and assembled in Paris — enough to validate the fundamentals in the lab, including the stability test.

The target is a factory floor, which sets the bar: a robot has to perform reliably for hours on end, and handle objects that are deformable, slippery, thin and fragile. The system retries when it makes a mistake and recovers from perturbations rather than failing outright. Being full stack is the point — co-developing the AI and the hardware in the same building allows fast iteration and hardware choices made to serve control and learning, and it means safety can be built in at every level rather than bolted on afterwards.

The approach follows from the research: scalable learning with minimal labelling, discovering capabilities bottom-up from cheap, continuous data collection rather than assigning tasks top-down and letting the data speak.

Safety — 2025

SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior

Pierre Sermanet, Anirudha Majumdar, Vikas Sindhwani

2025

Q: How would AI-powered robots behave if dropped into Science Fiction literature?
A: 95.8% aligned with humans (Sci-Fi decisions are only 21.2% aligned).

Q: Can we generate useful robot constitutions 📜 from Sci-Fi?
A: Sci-Fi inspired constitutions yield some of the strongest alignment on realistic safety benchmarks.

A benchmark built by mining the key moments in 824 major works of science fiction — films, television, novels and popular science books — where an AI or a robot made a decision that mattered. An LLM's recollection of each moment is used to generate a question about a comparable situation, the decision the fictional agent actually made, and the alternatives it could have chosen. Alignment is then measured against human-voted answers.

The first finding inverts the usual worry: modern LLMs paired with a constitution align with human values 95.8% of the time, while the decisions actually taken in science fiction align only 21.2% of the time. The genre is a catalogue of what not to do, and that turns out to be useful. Generated constitutions raise alignment over the base model from 79.4% to 95.8%, and hold up under adversarial prompting, where the base model collapses from 23.3% to a constitution-guided 92.3%. Those sci-fi-derived constitutions also score among the top performers on the ASIMOV benchmark, which is built from real hospital injury reports and real scenes — so rules learned from fiction transfer to reality. The released dataset holds 9,056 questions and 53,384 answers.

Safety — 2025

Generating Robot Constitutions & Benchmarks for Semantic Safety

Pierre Sermanet, Anirudha Majumdar, Alex Irpan, Dmitry Kalashnikov, Vikas Sindhwani

2025

Q: How can we ensure robots behave properly at scale?
A: Robot constitutions 📜!
Q: How do we verify behavior in undesirable situations at scale?
A: Generation!

We release the ASIMOV Benchmark for Semantic Safety of robots. We generate difficult-to-capture undesirable situations using image generation. We also generate robot constitutions straight from images.

Robotics safety used to mean collision avoidance. Once vision-language models started controlling robots capable of physical contact, with all their known failure modes, semantic safety became urgent: not whether the robot bumps into you, but whether it understands that what it is about to do is wrong.

The paper contributes a benchmark and a method. The benchmark is generated at scale rather than collected by hand — undesirable situations are synthesised from real visual scenes and from real hospital injury reports, using text and image generation to reach the difficult cases that are hard to capture in the wild. The method automatically generates robot constitutions from that data, with a novel auto-amending process that introduces nuance into written rules and measurably increases agreement with human judgments of what is safe and desirable. The paper explores the trade-off between general and specific rules across constitutions of very different lengths, reaching a top alignment rate of 84.3% — beating both no-constitution baselines and human-written constitutions. It deliberately declines to propose one universal constitution, arguing that rules need to be customised to legal, cultural and administrative context, and that a constitution's value lies in being human-readable and modifiable.

Reasoning — 2023

RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Pierre Sermanet, Tianli Ding, Jeffrey Zhao, Fei Xia, Yuan Cao et al.

ICRA 2024 · CoRL and NeurIPS 2023 workshops

Need help in the real world? RoboVQA can guide robots and humans through long-horizon tasks on a phone via Google Meet.

We propose scalable real-world data acquisition and augmentation strategies and release a dataset of 800k (video, question/answer) with robots & humans doing various long-horizon tasks.

We train a small video model (380M) that outperforms large state of the art zero-shot models by ~2x and demonstrate that scalable strategies for acquiring new data remain critical.

We use a speech intervention mechanism that automatically quantifies progress (intervention rate), provides corrections to retrain on, and allows performing tasks to completion in supervised real-world deployment.

A bottom-up data collection scheme that is 2.2× higher throughput than conventional top-down, step-by-step collection. Rather than scripting narrow tasks, operators fulfil any user request across three entire office buildings, using several embodiments: a robot, a human, and a human holding a grasping tool.

Models trained on all embodiments outperform models trained on robot data alone — even when evaluated purely on robot episodes — and for a fixed budget it pays to mix cheap human collection with expensive robot collection. The released dataset holds 829,502 video-text pairs across 29,520 unique instructions. A single video-conditioned model, RoboVQA-VideoCoCa, handles grounded high-level reasoning across these settings with a cognitive intervention rate 46% below the zero-shot state-of-the-art VLM baseline, and can guide real robots through long-horizon tasks. Video conditioning matters: video VLMs beat single-image VLMs by an average 19% error reduction across all task types. The intervention mechanism doubles as the evaluation metric and as the thing that makes imperfect systems deployable under human oversight.

Action · Robotics — 2023

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

A. Padalkar, et al.

CoRL 2023

A cross-institution effort pooling robot learning datasets across many embodiments, and the RT-X models trained on them.

Robotics conventionally trains a separate model per application, per robot, and often per environment. NLP and computer vision consolidated around general pretrained backbones years ago; this paper asks whether the same consolidation can happen for robot manipulation, and provides the data to find out.

Twenty-one institutions pooled demonstrations from 22 different robots into standardised formats — 527 skills across 160,266 tasks. Training a high-capacity model on that mixture produces RT-X, which exhibits positive transfer: individual robots get better by drawing on experience collected on entirely different platforms. The release is as much infrastructure as result, since the standardised dataset is what makes the question answerable by anyone else.

Action — 2023

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. Gonzalez Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, B. Zitkovich

CoRL 2023

Vision-language models trained on web-scale data are co-fine-tuned on robot trajectories, so that web knowledge transfers directly into closed-loop robotic control.

Vision-language models trained on internet-scale data are folded directly into end-to-end robot control. The recipe is deliberately simple: express robot actions as text tokens and put them into the training mixture exactly like natural language, then co-fine-tune a state-of-the-art VLM on both robot trajectories and web vision-language tasks such as visual question answering. Models built this way are what the paper names vision-language-action models.

Across roughly 6,000 evaluation trials, RT-2 shows capabilities that were never in the robot data. It generalises to novel objects, interprets commands absent from its trajectories — placing an object onto a particular number or icon — and does rudimentary reasoning, picking the smallest or largest object, or the one nearest another. Adding chain-of-thought reasoning extends this to multi-stage semantics: choosing a rock when asked for an improvised hammer, or an energy drink for someone who is tired. The web knowledge does not merely improve the policy; it changes what the policy can be asked to do.

Action · Reasoning — 2023

PaLM-E: An Embodied Multimodal Language Model

Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, Pete Florence

ICML 2023

An embodied multimodal language model that injects continuous sensor observations directly into the language model's embedding space.

Large language models reason well and understand nothing about the room they are in. PaLM-E attacks that grounding gap by pushing real-world continuous sensor data straight into the language model, rather than translating the world into text first.

Inputs are multimodal sentences that interleave visual, continuous state-estimation and textual encodings, trained end to end alongside a pretrained LLM across sequential manipulation planning, visual question answering and captioning. A single model handles varied observation modalities on multiple embodiments, and shows positive transfer — it benefits from joint training across internet-scale language, vision, and vision-language data rather than being diluted by it. The largest version, PaLM-E-562B, is simultaneously a robotics model and a visual-language generalist with state-of-the-art performance on OK-VQA, and retains its general language ability as it scales.

Reasoning — 2022

Inner Monologue: Embodied Reasoning through Planning with Language Models

Wenlong Huang*, Fei Xia*, Ted Xiao*, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, Brian Ichter

CoRL 2022

Language models plan robot behavior while continuously incorporating grounded feedback from the environment — an inner monologue that replans as the world answers back.

Planning with a language model in an embodied setting is not just a question of which skills to invoke, but how and when — and those answers change in response to the agent's own earlier choices. This work asks how far an LLM can get by reasoning over feedback expressed in natural language, with no additional training at all.

Feedback comes from several sources: success detection, scene description, and human interaction. Fed back into the prompt, they let the model form what the paper calls an inner monologue — a running, closed-loop commentary that it replans against. Closed-loop language feedback significantly improves completion of high-level instructions across three domains: simulated tabletop rearrangement, the same task on real hardware, and long-horizon mobile manipulation in a real kitchen. The result is notable for what it does not require: no fine-tuning, no new architecture, just the environment talking back in words the model already understands.

Reasoning · Action — 2022

SayCan: Do As I Can, Not As I Say — Grounding Language in Robotic Affordances

M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Fu, Ch. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. Jauregui Ruano, K. Jeffrey, S. Jesmonth, N. J Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, Cl. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, A. Zeng (alphabetically listed)

CoRL 2022

A language model proposes what to do; a learned affordance function scores what the robot can actually do here and now. The product of the two grounds language in the robot's real capabilities.

A language model asked how to clean a spill will produce a reasonable narrative that a particular robot in a particular kitchen may be entirely unable to carry out. SayCan closes that gap by pairing what the model wants to say with what the robot can actually do.

The language model supplies high-level semantic knowledge about how to accomplish an abstract, temporally extended instruction; a library of pretrained low-level skills supplies the grounding, with each skill's value function scoring how likely it is to succeed from the current state. Multiplying the two — what is useful to say against what is feasible to do — constrains the model to propose actions that are both sensible and possible. The robot becomes the language model's hands and eyes. Evaluated on real-world tasks with a mobile manipulator, the method completes long-horizon natural language instructions, and the ablations show that the real-world grounding is what makes the difference rather than the language model alone.

Action — 2020

Grounding Language in Play

Corey Lynch and Pierre Sermanet

RSS 2021

We present a simple and scalable approach for controlling robots with natural language: play through teleoperation, then answer “how do I go from start to finish?” for random episodes. We can then type in commands in real time.

By hooking up our English-trained model with a pre-trained language embedding trained on lots of text and different languages, it not only improves control but also allows commanding the model in 16 languages.

Combining natural language with play provides a breadth of skills while having no tasks determined in advance. This yields flexible specification of tasks — we can compose tasks on the fly: “pick up the object”, then “put the object in the trash”.

Imitation learning usually needs each task specified by a task ID or a goal image, neither of which is practical in an open world. Instruction-following work allows language, but typically assumes structure in the observations, the actuators or the language itself that does not survive contact with robotics.

This work learns perception from pixels, natural language understanding, and multitask continuous control end to end as one network — and, unusually, can absorb unlabelled, unstructured demonstration data carrying no task or language labels at all. That combination dramatically improves language-conditioned performance while cutting language annotation to under 1% of the data. At test time a single policy performs a wide variety of manipulation skills in a 3D environment, specified only by typed descriptions: "open the drawer… now pick up the block… now press the green button". Pairing the text-conditioned policy with a large pretrained language model makes it robust to out-of-distribution synonym instructions without collecting a single new demonstration, and extends it to commands in 16 languages.

Action · Latent Self-Supervision — 2019

Learning Latent Plans from Play

Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, Pierre Sermanet

CoRL 2019

How to scale up multi-task learning? Self-supervise plan representations from lots of cheap unlabeled play data — no RL was used.

Play is cheap and play is rich. It can be collected in quantity without segmenting tasks, labelling them, or resetting to an initial state, and it covers roughly four times more interaction space than task demonstrations gathered for the same amount of time. This paper turns that into a training signal.

Play-LMP self-supervises control on top of human teleoperated play, learning to organise play behaviours in a latent plan space and then reusing them at test time to reach specific goals. The shift is from a narrow, discrete set of tasks to the continuum of behaviours an environment actually affords. After self-supervising on unlabelled play alone, it substantially outperforms individual expert-trained policies on 18 difficult user-specified visual manipulation tasks. Two side effects are as interesting as the headline: play-supervised models are more robust to perturbation than their expert-trained counterparts, and they retry until they succeed. The agent also organises its latent plan space around functional tasks despite never having seen a task label.

Latent Self-Supervision — 2018

Self-Supervised Actionable Representations

Debidatta Dwibedi, Jonathan Tompson, Corey Lynch, Pierre Sermanet

IROS 2018

We learn continuous control entirely from raw pixels.

We use a multi-frame TCN to self-supervise task-agnostic representations from vision only, using 2 slightly different views of the cheetah. Then using RL on top of our embeddings we learn the cheetah task almost as well as if we were using the true proprioceptive states.

An extension of Time-Contrastive Networks that embeds several frames jointly rather than one frame at a time. A single frame can encode where things are; a short stack of them encodes how fast they are moving, and the paper shows this captures both position and velocity attributes substantially more accurately.

The test is whether such self-supervised representations are good enough to act on. Agents observing themselves taking random actions, or watching other agents perform tasks, learn embeddings that support continuous control policies trained with PPO using only those embeddings as input. On the real-world Pouring dataset the multi-frame model cuts error by 39.4% on motion attributes and 11.1% on static attributes relative to the single-frame baseline — the gap being largest exactly where motion matters.

Latent Self-Supervision · Robotics — 2017

Time-Contrastive Networks (TCN)

Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine

ICRA 2018

We propose a general self-supervised method for learning representations from raw unlabeled videos, and show that the representations are rich enough to perform robotic tasks.

We use the distance in our learned embedding space to a video demonstration as a reward. An RL algorithm can learn to perform a pouring task using this reward: the robot learned to pour in only 9 iterations from a single video demonstration, while never receiving any labels.

We also show that a robot can teach itself how to imitate people. By training a single TCN on videos of both humans and robots performing random motions, the model finds correspondences between humans and robots despite never being given any label correspondences.

A self-supervised approach that learns representations and robot behaviours entirely from unlabelled video recorded from multiple viewpoints. Imitating human behaviour requires a viewpoint-invariant representation capturing the relationship between end-effectors and the environment, object attributes, and body pose — none of which is available as a label.

The training signal is a metric loss over time and viewpoint: simultaneous views of the same instant are attracted in the embedding space, while temporal neighbours — which look similar but are functionally different — are repelled. The model therefore learns simultaneously what is common between different-looking images and what differs between similar-looking ones, discovering attributes that hold across viewpoint but change over time, and ignoring nuisance variables like occlusion, motion blur, lighting and background. The representation supports two robot imitation settings: mimicking human poses with no explicit correspondence between human and robot bodies, and serving as a reward function inside a reinforcement learning loop, where the robot learns to pour from a single video demonstration in nine iterations without ever receiving a label.

Latent Self-Supervision · Robotics — 2017

Unsupervised Perceptual Rewards

Pierre Sermanet, Kelvin Xu, Sergey Levine

RSS 2017

We propose learning unsupervised perceptual rewards that can be fed to an RL system, and show it is able to learn a robotic task such as door opening from a few human demonstrations.

Reward design and exploration time are arguably the two biggest obstacles to deploying reinforcement learning in the real world. Designing a reward often means hand engineering, and sometimes means bolting extra sensors onto the environment purely to detect whether the task succeeded. Worse, interesting tasks have implicit intermediate steps that a final-outcome reward says nothing about.

This work infers perceptual reward functions from a handful of demonstrations by exploiting the abstractions already present in a pretrained deep model. The method identifies a task's key intermediate steps from only a few demonstration sequences and automatically selects the features most discriminative for recognising each step, with no explicit specification of sub-goals. The resulting rewards are dense and smooth — the properties an RL agent actually needs — and are shown to drive learning of a real-world task such as door opening from a small number of human demonstrations.

Vision · Deep Learning — 2015

Visual Attention

Pierre Sermanet, Andrea Frome, Esteban Real

ICLR 2015 (workshop)

We demonstrate a foveated attention RNN that is able to perform fine-grained classification. Tracking naturally emerges from our foveated model when run on videos, even though it was only trained on still images.

An extension of recurrent attention models into a visually unconstrained setting: fine-grained categorisation on the Stanford Dogs dataset, with real clutter, occlusion, lighting and pose variation. Prior attention work had largely stayed on toy or constrained problems such as MNIST digits.

The RNN structure is retained, but the visual network is far more powerful and is pre-trained at large scale outside the attention loop. The model learns to direct high-resolution attention to the most discriminative regions with no spatial supervision — no bounding boxes, no part annotations — and discriminates dog breeds moderately well even when given only a low-resolution context image plus a few narrow, cheap glimpses at faces and fur patterns. It outperforms the GoogLeNet classification model on this task. The broader argument is architectural: attention models train end to end, unlike detection pipelines assembled from hand-engineered stages that discard information at every seam. Run on video, tracking emerges from a model trained only on stills.

Vision · Deep Learning — 2014

Inception / GoogLeNet

Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich

CVPR 2015 · CVPR 2025 ten-year impact prize

A deep architecture for computer vision. Our model obtained 1st place for the classification and detection tasks in the 2014 ImageNet Challenge.

The architecture that won ImageNet 2014 classification and detection, and the reasoning behind it. The paper's stated concern is not accuracy alone but efficiency: with mobile and embedded computing gaining ground, power and memory use matter, so most experiments were held to a budget of roughly 1.5 billion multiply-adds at inference — a deliberate constraint intended to keep the result usable rather than academic.

"Deep" is meant in two senses: greater network depth, and a new level of organisation in the form of the Inception module, which computes 1×1, 3×3 and 5×5 convolutions plus pooling in parallel and concatenates them, letting the network choose its own filter scale at every stage rather than having it fixed by the designer. The dimensionality-reduction variant inserts 1×1 convolutions before the expensive filters, which is what keeps the computational budget in reach. The paper also credits the era's progress to the synergy of deep architectures with classical computer vision rather than to depth alone.

Vision · Deep Learning — 2014

Dogs vs. Cats Kaggle challenge

Pierre Sermanet

2014

1st place in an image classification Kaggle challenge between dog and cat images. Most of the top entries are based on our OverFeat model.

Vision · Deep Learning — 2013

OverFeat

Pierre Sermanet, David Eigen, Xiang Zhang, Michael Mathieu, Rob Fergus, Yann LeCun

ICLR 2014

This model obtained 1st place in the 2013 ImageNet object localization challenge. The model and pre-trained features were later released to the public.

OverFeat has been used by Apple for on-device face detection in iPhones: blogpost.

An integrated framework using a single ConvNet for classification, localization and detection at once. The central observation is that a multiscale sliding-window approach can be implemented efficiently inside a ConvNet itself, since the convolutional layers already share computation across overlapping windows — making dense evaluation affordable rather than prohibitive.

To this it adds a deep learning approach to localization that predicts object boundaries directly, and an unusual choice at inference: bounding boxes are accumulated rather than suppressed, so agreement between overlapping predictions raises detection confidence instead of discarding evidence. All three tasks are learned simultaneously in one shared network. The framework won the localization task of ILSVRC 2013 and was competitive on detection and classification; post-competition work set a new state of the art on detection. The released feature extractor, OverFeat, became one of the first widely used pretrained deep vision models — later used by Apple for on-device face detection in iPhones.

Vision · Deep Learning — 2013

Pedestrian Detection

Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala, Yann LeCun

CVPR 2013

State of the art results on pedestrian detection datasets using deep ConvNets in the EBLearn framework.

State-of-the-art and competitive results across all the major pedestrian detection datasets with a single convolutional network, at a point when the field was still dominated by hand-crafted features.

Three departures from a standard ConvNet carry the result. Multi-stage features, with connections that skip layers, let global shape information be integrated with local distinctive motifs rather than the classifier seeing only the most abstract stage. And the filters at each stage are pre-trained with an unsupervised method based on convolutional sparse coding — the learned first-layer filters include not only oriented edge detectors but corner and junction detectors, features that would be tedious to design by hand. The system was built in the EBLearn framework.

Vision · Deep Learning — 2012

Convolutional Neural Networks Applied to House Numbers Digit Classification

Pierre Sermanet, Soumith Chintala, Yann LeCun

ICPR 2012

State of the art results in house numbers classification using deep ConvNets.

Classifying digits from real-world house numbers — the SVHN dataset — with convolutional networks, at a time when most vision systems still relied on hand-designed features. The argument is that a ConvNet learns a feature set optimised for the task at hand instead of one chosen in advance.

Two departures from the standard architecture do the work. Multi-stage features feed the classifier the outputs of earlier stages as well as the last one, so local detail reaches the decision alongside global shape. And Lp pooling replaces the usual average or max pooling, with the paper analysing how the choice of p affects validation error. Together they establish a new state of the art of 94.85% on SVHN, a 45.2% relative error improvement. The source code and a tutorial were released through EBLearn.

Vision · Deep Learning — 2011

Traffic Sign Recognition

Pierre Sermanet, Yann LeCun

IJCNN 2011

This deep model obtained 2nd place in a traffic sign recognition challenge using the EBLearn framework. It uses skip connections in deep ConvNets to better combine low-level and high-level learned features.

ConvNets applied to the GTSRB German traffic sign benchmark. Signs are rigid and designed to be visible, which makes the task sound constrained, but the dataset is full of real-world damage: viewpoint variation, saturation and low contrast, motion blur, occlusion, sun glare, faded colours, graffiti, stickers, and inputs as small as 15×15 pixels.

The architecture departs from a standard ConvNet by feeding the classifier the pooled first-stage features alongside the second-stage ones. The second stage carries global, invariant shape; the first carries local motifs with more precise detail; together they give the classifier several scales of receptive field. It also replaces the usual pointwise sigmoid with a rectified sigmoid followed by subtractive and divisive local normalisation, and trains on a jittered set five times the original size, perturbed in position, scale and rotation. The system placed second in phase I of the competition at 98.97%, above the 98.81% human benchmark and a hair behind the winner. Further experiments — a deeper classifier and, unexpectedly, dropping colour for grayscale — set a new record of 99.17%. Randomly-initialised features alone still reached 97.33%, which the paper uses as a cheap way to search architectures before committing to training.

Latent Self-Supervision · Vision · Deep Learning — 2010

Unsupervised Convolutional Feature Hierarchies

Koray Kavukcuoglu, Pierre Sermanet, Y-Lan Boureau, Karol Gregor, Michael Mathieu, Yann LeCun

NIPS 2010

An unsupervised method for learning multi-stage hierarchies of sparse convolutional features. One of the few instances of this period where unsupervised pretraining improved results in a supervised task.

Sparse coding had become a popular way to learn visual features, but it was almost always trained on isolated patches. Applying the resulting filters convolutionally then produces highly redundant codes, because every overlapping patch was encoded without reference to its neighbours. This paper trains convolutionally over large image windows instead, which removes that redundancy between neighbouring feature vectors and makes the representation more efficient.

Alongside the linear decoder that reconstructs the image from sparse features, the method trains a feed-forward encoder that predicts quasi-sparse features directly from the input, so inference at test time costs one pass rather than an optimisation. The change in training regime changes what gets learned: patch-based training rarely yields anything but oriented edge detectors, while convolutional training produces markedly more diverse filters — centre-surround, corner detectors, cross detectors and oriented gratings. Stacked into a multi-stage convolutional architecture — 64 features at the first stage, 256 at the second, connected through a sparse table amounting to 4,096 kernels, more dictionary elements than prior convolutional RBM or sparse coding work had attempted — these filters improve results across several recognition and detection tasks, including Caltech-101 and INRIA pedestrians.

Vision · Deep Learning — 2009

EBLearn

Pierre Sermanet, Koray Kavukcuoglu, Yann LeCun — additional help from Soumith Chintala

ICTAI 2009

A C++ deep learning framework similar to Torch and used for multiple state of the art results in computer vision.

Energy-based learning offers a single framework for supervised and unsupervised training of probabilistic and non-probabilistic factor graphs. A model assigns a scalar energy to configurations of inputs, outputs and latent variables; inference finds the configuration that minimises that energy; learning shapes the energy surface so that correct outputs sit lower than all incorrect ones.

EBLearn is an open-source, cross-platform C++ library for building such models. It has two components: libidx, an efficient and flexible multi-dimensional tensor library, and libeblearn, an object-oriented library of trainable modules and learning algorithms with support for convolutional networks, image processing and graphical display. Machines are assembled from predefined modules and loss functions, with gradient-based learning implemented through semi-automatic differentiation of the assembled model. The paper works through how regression, classification, unsupervised algorithms and full vision architectures such as LeNet-7 all map onto the same formulation, and shows the library detecting objects with a trained detector.

Robotics — 2009

Teaching Assistant for NYU Robotics class

Pierre Sermanet, Yann LeCun

2009

Teaching assistant for the NYU robotics course, building small mobile robot platforms for student projects.

Robotics — 2005

LAGR: Learning Applied to Ground Robots

Yann LeCun, Urs Muller, Pierre Sermanet, Marco Scoffier, Chris Crudelle, Beat Flepp, Ayse Erkan, Matt Grimes, Raia Hadsell, Koray Kavukcuoglu, Marc'Aurelio Ranzato, Jan Ben, Sumit Chopra, Jeff Han, Marc Peyote, Ilya Rosenberg, Yury Sulsky

DARPA, 2004–2008

A DARPA challenge where the NYU-NetScale team developed ConvNets for long-range off-road navigation from 2004 to 2008.

Robotics · Vision — 2009

Learning Long-Range Vision for Autonomous Off-Road Driving

Raia Hadsell, Pierre Sermanet, Jan Ben, Ayse Erkan, Marco Scoffier, Koray Kavukcuoglu, Urs Muller, Yann LeCun

JFR 2009

An overview paper of our self-supervised deep learning vision model.

The journal-length account of the self-supervised long-range vision system. A deep hierarchical network extracts features from camera images; a realtime classifier trained on those features predicts traversability from 5 metres to beyond 100, far past the 12-metre limit of the stereo module that supervises it.

The paper credits the result to the quality of the training data generated on every frame: robust, visually consistent stereo labels; wide-context input windows normalised so objects appear at consistent scale regardless of distance; and a concise, discriminative feature representation. It covers the feature learning in detail — including unsupervised convolutional auto-encoder training and radial basis function features — along with the five-category labelling scheme, ground plane extraction, and the failure modes the labelling has to survive. Evaluation uses both a ground-truth dataset and field runs, where robots with long-range learning held to paths that the short-range system repeatedly lost, needing rescue from the undergrowth.

Robotics — 2009

Collision-Free Off-Road Robot Navigation

Pierre Sermanet, Raia Hadsell, Marco Scoffier, Matt Grimes, Jan Ben, Ayse Erkan, Chris Crudele, Urs Muller, Yann LeCun

JFR 2009

An overview paper of our navigation system, designed to naturally handle errors and outputs coming out of a deep vision model. This model decouples the fast and short-range navigation from the slow and long-range navigation to achieve robustness.

The complete navigation architecture behind the LAGR system, and the argument for splitting it by range. Humans combine cues at several spatial resolutions and time scales: distant gazing sets the general direction and is refined slowly, while nearby obstacles demand constant attention and fast reaction. The system mirrors that.

A short-range perception loop runs at high frame rate on low-resolution images and handles local planning and obstacle avoidance; a long-range adaptive vision loop runs at lower frame rate on maximum-resolution images and handles strategic planning. Probabilistic traversability labels from both accumulate into a robot-centred hyperbolic-polar map with a 200-metre effective range. Short-range planning uses the recorded maneuver dictionary rather than a dynamical model. Localisation combines GPS, wheel odometry, an IMU and a fast, low-complexity rotational visual odometry module. The layers are deliberately independent, so each is insulated from the latency of the more complex ones above it: wheel commands at 20Hz, reactive planning at 5–10Hz, deliberative planning at 1Hz. The system beat both the reference baseline and its own configuration with long-range vision disabled, and was verified in independent government tests.

Robotics — 2008

Learning Maneuver Dictionaries for Ground Robot Planning

Pierre Sermanet, Marco Scoffier, Chris Crudele, Urs Muller, Yann LeCun

ISR 2008

Instead of computing the theoretical dynamics of a vehicle, we propose to simply record the observed dynamics while a human operator “plays” with the robot, essentially trying all possible moves. At test time, the model has a bank of observed possible trajectories for every state of the motors. Trajectories leading to collisions are discarded, while the fastest available trajectory is selected. While we observed many collisions using the baseline system, we did not observe collisions after introducing this model.

Vehicle dynamics are normally captured by a model whose parameters come from system identification or from the vehicle's specifications. Those models are accurate in theory but blind to differences between individual vehicles, awkward to adapt to new environments, and often need hand-crafted heuristics for cases like reversing. This work replaces the model with a recording.

While a human drives — or the robot drives itself — every traversed trajectory is stored in a bank indexed by the initial left and right wheel speeds and by where it ends on a 2.5m circle. Ten speed bins per wheel give 100 states, each holding around 100 trajectories reaching different angles. At planning time the current wheel speeds select the feasible set, trajectories crossing non-traversable cells are discarded, and the cheapest survivor is executed as a list of wheel commands, at almost no online computation cost. A bank only 15% full, extracted from two hours of human driving, already drove well; 18 hours of logs filled it to 64% and covered a wider range of situations. Collisions were frequent with the baseline system and were not observed once this one was in place.

Robotics — 2008

Mapping and Planning under Uncertainty in Mobile Robots with Long-Range Perception

Pierre Sermanet, Raia Hadsell, Marco Scoffier, Urs Muller, Yann LeCun

IROS 2008

A hyperbolic-polar coordinate mapping system that is naturally suited to handle imprecisions in long-range visual navigation.

Self-supervised long-range vision can label obstacles and pathways a hundred metres out, but both the category and the range of distant regions carry considerable uncertainty. This paper builds a map that represents both, and accumulates evidence across frames rather than overwriting it.

A pixel covers a near-constant angular extent but a wildly varying span of distances — a few centimetres near the robot, effectively infinite near the horizon. The map geometry follows that: hyperbolic-polar, with constant 20cm radial resolution out to about 15m, then hyperbolically growing cells to 200m and beyond, so an effectively infinite radius fits in a finite grid. A robot-centred Cartesian map at the same resolution would have needed 500×500 cells. Every cell holds a histogram accumulating per-class evidence from the classifier's five categories, so the traversability decision is deferred to planning time and the policy can be made more conservative or more aggressive without rebuilding the map. Because the map is fixed to the local pose frame, recentring as the robot moves costs only a translation.

Robotics · Latent Self-Supervision · Deep Learning — 2008

Deep Belief Net Learning in a Long-Range Vision System

Raia Hadsell, Ayse Erkan, Pierre Sermanet, Marco Scoffier, Urs Muller, Yann LeCun

IROS 2008

Self-supervised long-range visual navigation with deep ConvNets.

A deep belief network extracts features from the robot's camera images, and those features train a realtime classifier that predicts traversability all the way to the horizon, enabling strategic rather than merely reactive planning.

Three choices make the self-supervision work. Training windows are large enough to carry context, not just colour and texture. The stereo supervisor emits five visually distinct categories — super-ground, ground, footline, obstacle, super-obstacle — chosen so the labels stay consistent instead of noisy, since inconsistent supervision causes the learning to fail outright. And a horizon-levelled input pyramid normalises for distance, so similar obstacles appear at similar heights regardless of how far away they are. The classifier's output populates a hyperbolic-polar costmap on which the planner runs. The system was developed and tested on the DARPA LAGR platform, both standalone and inside the full navigation stack, where long-range vision produced measurably smoother and more far-sighted driving.

Robotics · Latent Self-Supervision — 2007

Online Learning for Offroad Robots

Raia Hadsell, Pierre Sermanet, Ayse Naz Erkan, Jan Ben, Jefferson Han, Beat Flepp, Urs Muller, Yann LeCun

RSS 2007

Online adaptation of long-range vision by self-supervising with short-range stereo vision.

Stereo range estimates become unreliable past 10 to 12 metres, which leaves a robot driving in a self-imposed fog — entering dead ends and slowly discovering pathways a human would see immediately. This system trains a classifier online, on every frame, from the stereo module's sparse near-range traversability labels, then applies it to the entire scene.

Two ideas carry it. A distance-normalised image pyramid makes it practical to train on large, context-rich windows rather than colour and texture patches. And spatial label propagation: once stereo labels a location, that label is propagated back to every earlier view of the same target, so the classifier learns view-invariant classifications and can train on far-away views captured before any label was available. A ring buffer acts as short-term memory, keeping the discriminative training balanced and consistent. The result sees obstacles and paths at 30 to 40 metres, far beyond stereo range, and adapts quickly when the environment changes.

Robotics · Vision — 2004

EUROBOT 2004 Competition

Computer vision, navigation and behaviors by Pierre Sermanet, Philippe Rambert, Jean-Baptiste Mouret. Entire team: Evolutek

2004

Vision-based behaviors in a robot-rugby challenge.