The talk given by Georgia Chalvatzaki in KUIS AI Talks on November 28, 2023.

Title: Interactive Robot Perception and Learning for Mobile Manipulation

Abstract:
The long-standing ambition for autonomous, intelligent service robots that are seamlessly integrated into our everyday environments is yet to become a reality. Humans develop comprehension of their embodiments by interpreting their actions within the world and acting reciprocally to perceive it —- the environment affects our actions, and our actions simultaneously affect our environment. Besides great advances in robotics and Artificial Intelligence (AI), e.g., through better hardware designs or algorithms incorporating advances in Deep Learning in robotics, we are still far from achieving robotic embodied intelligence. The challenge of attaining artificial embodied intelligence — intelligence that originates and evolves through an agent’s sensorimotor interaction with its environment — is a topic of substantial scientific investigation and is still an open challenge. In this talk, I will walk you through our recent research works for endowing robots with spatial intelligence through perception and interaction to coordinate and acquire skills that are necessary for their promising real-world applications. In particular, we will see how we can use robotic priors for learning to coordinate mobile manipulation robots, how neural representations can allow for learning policies and safe interactions, and, at the crux, how we can leverage those representations to allow the robot to understand and interact with a scene, or guide it to acquire more “information” while acting in a task-oriented manner.

Short Bio:
Georgia Chalvatzaki, recently promoted to Full Professor of Interactive Robot Perception and Learning in April 2023, holds a joint appointment at the Computer Science Department of the Technical University of Darmstadt and Hessian.AI. Prior to her promotion, she served as an Assistant Professor and Independent Research Group Leader, having secured the prestigious Emmy Noether grant from the German Research Foundation (DFG) in March 2021. She completed her Ph.D. in 2019 at the National Technical University of Athens, Greece, where she was a part of the Intelligent Robotics and Automation Lab within the Electrical and Computer Engineering School. Her doctoral thesis, titled “Human-Centered Modeling for Assistive Robotics: Stochastic Estimation and Robot Learning in Decision-Making,” laid the foundation for her current research interests, which include robot learning, planning, and perception.

Everyone today we have Georgia and Georgia is a full professor at as of late last April and also affiliated with head cni she was an assistant professor there and independent research group leader um she has secured the prestigious em notary Grant from German Research Foundation DFG um in March 21 and she

Received her PhD from Technical University of ethens in Greece um so today she’s going to talk about interactive robot perception and learning for mobile manipulation we are looking forward to her talk thank you very much Storia stage is yours thank you thank you very much for inviting me

So um today I will present to you a subset of works from uh from my group that targets mobile manipulation that it is uh a problem that I’m very passionate about I think that it is the the next thing that we should focus on in in

Robot learning and uh I will try to to convince you uh about this and why we we should put our effort so if you work in robotics in robotic manipulation I think we should put our efforts together to um bridge the gap between navigation H manipulation and of course perception um

Okay so if we if we look at at at humans even at very very early age um we would see that um uh the kids uh usually interact with their environment they try to uh to understand how the world uh works right how their own Body Works um

And basically what we see is this internal Loop of uh perceiving in order to act and act in order to uh to P to perceive and by this uh by this cycle uh we as humans very early on we develop sensory motor coordination so we uh basically do also this uh handai

Coordination we understand our own body uh what we are capable of we develop skills and we create a representation of the world so we create this external representation of the world in which we uh which we act and uh finally we interactively explore and learn new skills and this is a continuous

Um uh cycle that happens in in um a human’s uh life um now the question is why don’t we have this still in robotics that the the the the points are many right so uh for sure one part comes from the lack of uh good perception

Algorithms and of course then how we can couple them with robotics but obviously also the hardware was not good enough uh in in the earlier uh days and now I believe that we have an opportunity to um to develop algorithms to towards embodied uh uh artificial intelligence through mobile manipulation uh systems

In general and as I said from the beginning why I like mobile manipulation systems well they combine mobility and manipulability and they are equipped with multiple sensors so that they can uh sense their world around them they have an increased workpace so in the whole world is their workspace so uh

They could reach possibly um every space that a human can do I have to note here that we could even frame uh humanoid robots as mobile manipulation robots on every type of mobile Bas where a manipulator can move is a mobile manipulation system so even an aerial

Robot with an arm is a mobile manipulation system so we can uh now Envision all the applications of this uh types of robot but of course the uncertainty and the difficulties that arise from combining manipulation with Mobility are much much larger um and what I propos in my research in order to

Uh uh basically unlock the path to uh to robotic and boded intelligence is to leverage structure and uh in this talk I will focus on on three topics first on how we can leverage structure in order to learn uh mobile manipulation coordination uh how we can extract representations uh for mobile

Manipulation robots in order to interact with their environment and finally how they can how we can enable them to actively explore and act in open SC in house let’s let’s say hous like U environments um so uh the first part of of the talk so I selected three topics

So that I can uh maybe keep it short and we can have like longer uh discussion um the first part um targets the coordination of mobile manipulator so um as I said before mobile manipulators combine all these different embodiments right so if we see this uh robot that we have in our lab

The Thiago robot uh it has two uh humanoid arms s degrees of freedom itats it has a a a torso um it has a two degrees of freedom uh head and a mobile base in this particular case this robot has a differential Drive uh mobile base

And it is equipped also with cameras and Laser sensors and other types of sensors now um we can either try to solve this uh problem of mobile manipulation as a whole thing uh but this becomes very difficult because you have a hyper redundant uh robot so you have like a

20° Freedom robot when you try to just grasp an object that it is a six dimensional po so you can understand that there is a big gap so you have a very hyper robot but you have too many constraints this is the problem you have too many constraints especially when you

Are acting in the real reward then you want to um resolve not only collisions but also for example time uh efficiency in energy all these are uh basically additional constraints that we need to add um into our optimization ideally so this makes it quite challenging to solve

Uh whole body control uh and whole body planning uh problems um so in in a work that we presented last year in in iros it was a r paper that was also presented in iros and received the best paper award in Mobile manipulation we um consider to decouple uh this problem and

Consider uh basically the uh the robot as a two agent system in this particular case so we didn’t tackle all the embodiment of the robot in this world so we consider just the base and uh the left arm of the robot and we Tred to coordinate them and let me explain uh

What I mean by coordination and what I I mean by a hybrid action space reinforcement learning that you see uh here so in when a human moves around right uh we move around and we want to uh go for example in this white table and grasp maybe the yellow sugar book uh

The human knows uh what is its own reach what is uh their limit so they know that they have to move close to the table and once they are close to the table they can now extend the arm and uh basically execute a manipulation action we want to uh basically Endo the mobile

Manipulation robots with the same capability so the robot to understand its own reach and understand to take decisions uh according to when it can really manipulate an object or not uh therefore we create we we uh formalized a a discrete action space uh where the the discrete decision sorry a hybrid

Action space with a discrete part uh that is when should I activate my arm in order to execute a manipulation action and the continuous part that is the optimal of the robot that it is expressed in SC2 based on the uh pose of the base so U the position and

Orientation of the base that also dictates the whole body uh Poe and it also correlates with theability of the robotic arm so if the robot for example position itself with its back to the table it is evident that okay the the solution space from uh a planning um

Algorithm or an inverse kinematic um uh solver might be difficult to reach the sugar box in this uh position right um so how we uh try to formalize this uh problem uh we try to use it in reinforcement learning so that we can reward the agent whenever it positions

And execut a manipulation action uh correctly and we uh how we deployed this um uh hybrid action phas we uh uh basically Define this hybrid uh policy that contains two components uh where the uh discrete uh part that you see here the uh Peta D um conditions on

The continuous actions that we take and intuitively if we think about the way that we Define here the action space the agent has to First Position itself so the continuous uh take the continuous action and then based on this continuous action the robot should decide whether

To activate uh the arm or not uh and we can train this uh in with our favorite deep reinforcement learning algorithm by also uh deploying a um par reparameterization uh trick so we can uh model the discrete variable a categorical uh distribution uh in particular here it’s a 01 but it could

Have been any type of um uh discrete uh distribution and you can deploy the Gumble Toof Max reparameterization in order to formalize this categorical uh distribution and take gradients and do back propagation uh in order to alter the parameters of this uh policy okay so now we formalize the uh the hybrid

Policy for this uh mobile manipulation uh scenario and now we need to train it so now we need to collect the data in order to ensure that the the robot will understand its own reach uh right so it’s it’s basically understanding uh if I am studying here

Can I reach this object to reach it and select this kinematically reachable uh poses now in classical robotics how uh this has been uh tackled in in in the past um uh people would uh basically take the operational workspace of the robot so what you uh see uh here is for

The left arm of of the robot it’s the operational workspace uh this is quite easy to uh to extract so you put your robot in a simulator you densely um sample uh 6D poses in in uh ellipses or yeah uh depending of course of the of

The structure of the robot here you can see basically the disability uh the workspace of the robot so you sample the workspace of the robot and then you uh try to find how many ik solutions you can have with what frequency you can find an actual solution given random

Initializations of the ik uh Problem Solver so you create this kind of frequency uh map uh about your workspace and what we can see with Bluer are basically the areas that are better reachable for this type of robot and you can see also the reability on on laterally so um now because we’re

Considering mobile manipulation robots in order to understand how to place the robot in order to reach reachable uh Target um we consider that the floating base scenario so we consider that the the base can be move everywhere in in the space and we invert the mapping of this operational workspace so we invert

This map and then we basically slice it on the level of the floor and by slicing it on the level of the floor we get what we called the inverse reability uh map a note here is that in in this reability map and in the original sorry on the original um uh

Workpace we can also include the measure of manipulability so not only the I solution but also add metric of manipulability that shows us also for manipulation how easy it is to continue manipulating an object if we place a robot in a particular location so with this inverse map we know that if we

Place our robot in the Bluer areas uh there is a chance of high man ability to um find a pose that would make our robot Reach This 60 Target that you see here with purpose um what this means is that the that the robot will have to uh query

For any different 60 Target we will have to query a different map so with every 60 Target we get a different uh mapping on the floor on how we can place the robot so usually what how this was done people would um randomly sample 60 Target and then store them in a memory

And if of course you had to reach a Target that did not exist in your memory you have to basically requery and do I Solutions and so on if you had obstacles around you would have to cut out so you have to do several charistics in order

To cut out uh the map so and those another problem is that this is quite discrete right um this is not a continuous uh representation so we thought like why don’t we use this information in order to train um a reinforcement learning agent so to learn

A cur ability map uh so we use this uh data that we collected anyways by random sampling and querying and so on and train a q function on on this base placement uh problem and uh not only we can consider now uh the uh manipulability so the overall reability

But the actual task success so the the uh reward function for this problem contains actually the execution of the uh reaching uh action and you can see a completely different curent uh map uh so the the robot has many more capabilities in the end in actually placing itself

And reaching the end goal rather than just discretizing and considering the overall uh manipulability and we get this smooth cability map so now we have this local map that it does it just tells us what is the next step that the robot should do what is the next poll that the robot should

Take in order to reach this target but this is very very local and mobile manipulation is inherently long Horizon the robot has to navigate an environment and find the targets manipulate them and so on so we need to somehow deploy this knowledge about theability into a framework that can solve longer Horizon

Problems H and use it in essence as a prior basically and uh in order to use this prior uh knowledge we uh proposed boosted hybrid r enforcement learning that it is an algorithm that allows us to transfer Knowledge from Behavior priors and build more complex uh

Behavior so uh with uh uh boosting uh in reinforcement learning that has been uh uh some uh prior works you can find the references here uh what we can do is that we can uh start from a Bas uh problem where we uh formalize as a reinforcement learning problem and we

Solve with and we learn uh sorry A Q function we can then progressively transfer this uh Q function onto the next problem that it is of higher difficulty in theory and don’t train for a new Q function but rather keep the prior knowledge fixed so the prior to

Function uh fixed and learn only a residual over it uh there therefore you only learn what you are missing on representing well what are good or bad actions to take for solving a more challenging problem and you can keep building on this so you can keep building uh Q functions that are more

Complex by adding up uh the uh previously learned component uh so uh the original function and the residuals of these tasks that are way more challenging um um up to a point of course that you uh that that you stop in theory theoretically you can add up as

Many as you want and you can have the exact same representation of the original function but what is now um favorable in deep reinforcement learning is that usually when you have tasks that are uh very difficult and you might have complex reward functions it might be very difficult with a single function

Approximator with a single neural network to present everything so by uh modularizing and basically splitting down the problem into this um uh residuals you make your life easier you make the learning uh easier and you also can deploy already existing knowledge so if you have learned already where when

Should I where should I position myself in order to activate my arm and do a reaching action why should it be learned again when you want to for example uh grasp an object or navigate and then avoid obstacle if this has been learned already um in order to also facilitate

The training of the of the critic we have seen that it is uh quite uh beneficial to also update the actor uh the actor is not transferred so the actor is learned from scratch uh over these composed uh two function composed of these uh res ual but what was very

Beneficial is to add an additional KL regularization uh component where you reward in essence the agent for remembering actions that were good from your previous uh task and this is why you see here this uh forward KL regularization um what is also very important in this uh framework that handles uh

That that basically represent the Q function as this sum of residuals and that it doesn’t really create residuals on the policy space but rather on the Q function uh space is that we can handle uh the problems of changing reward functions across that and we can also handle the changing observation spaces

Across that so if we keep building on on prior knowledge that was simpler and like initially you were just trying to understand how your body works and then once you capitalized on that you can add on on manipulating object the observation space increases because now it’s not just you just you and the

Objects around you uh so your observation of your State uh augments and this can be easily handled this asymmetric information while transferring to uh Tas of higher complexity complexity can be nicely handled by uh um this framework because the state can effectively add on uh new information while your residual can just

Keep looking at the F the part of the state that was relevant to to each own uh part uh while it was trained and the same goes for the policy uh the prior policy can just take the part of the observation that was relevant for this particular policy and you just your new

Policy on the new observation bace so we started by just doing this uh reability initially in this locality the one step basically motion and uh activating the arm and solving for the reaching task having to navigate then having to avoid additionally obstacles then having different of types of configurations

Trying to um grasp an object not only reach re and different types of configurations and also when there are other objects on the table that uh create clut and you have to find basically the clearance in order to uh grasp an object and what is very interesting is that uh this approach can

Directly be transferred uh in the real world because in essence we are planning we are deciding we are doing learning on a level higher than the actual execution so we decide on poses of the robot and whether it should activate the arm or not and on the low level in order to

Execute this you can deploy any controller or planning algorithm in order to resolve this and you can also deploy this sequentially in order to do sequential placing and rearrangement task because the agent if it knows its own reach it is able to solve uh grasping placing and other types of uh

Problems and currently we are extending this work to also consider real uh contact based manipulation because you may keep requiring the high manipulability regions in order to keep um manipulating an object one one caveat here was that if you train of course in simulation with a naive uh navigation

Policy for example uh uh this comes at a cost so it increases the S to real Gap so we had to uh basically tune our navigation policy to resemble the one that we used uh in simulation but this also can be resolved if you use the same

Exact strategy uh that that you will use in the real uh world for um moving the robot around in in the room is there any question of far if if there is if there is no okay I will continue to the next uh uh part I

Just wanted to to to give you just a bit of intuition about also uh when when we want to do this whole body control basically coordination that I was describing in the beginning of uh of the talk we may even assume that we have some skills that we have learned that we have

Designed for example and uh if we look in into classical robotics again we will see two types of ways in order to control our uh robot the one is this reactive robot motion controllers uh that uh usually when we design them as like operational space controller they

Are um computationally um uh cheap uh they uh they are having high frequencies H but the drawback is that they are quite myopic so they don’t have this look Ahad they cannot predict how the environment might change so that they can arrange the strategies accordingly

And on the other hand we have this uh planning based or look ahead motion generation uh strategies where we can provide feasible uh trajectories with higher success guarantees but the drop is higher computational load and low control frequency cont compared to these uh reactive uh motion controllers um so

Very briefly I will go through this uh because I think it’s inter might be interesting to some people so if we uh consider that our skills can be represented as um uh basically uh energy based policies or uh in general that they belong in the exponential families

Uh so uh that they belong they are a distribution that belongs to the exponential family then it is very easy to compose uh this kind of skill uh right because we can just uh uh add them up uh the components of the of the exponent and create what we call a

Product of experts basically which can be represented as a policy that is composed by other prior uh skills and which is weighted by a temperature parameter or a weight in General that um uh handles how much contribution each expert has in these uh SE position and

Then if you want to querry an optimal action you can just do uh uh maximum likelihood and exact the action that uh uh basically uh maximizes the uh the policy uh one caveat here is how do you choose this temperature V or waiting parav vaa according to the way that your

Environment changes and for that we have proposed two different uh ways of handling this problem the one is probabilistic and the other one is basically optimization based uh based on optimal uh transport so if we frame the problem of selecting the best possible uh vaa parameters for adjusting the uh the

Importance of our expert according to uh the environment or according to the time step that we are in the environment um um we can very easily frame it as a planning and difference approach and plan on the parameter uh space but then what is also interesting to see is what

Happens if we formalize this problem as an optimal transport problem where we consider that every skill is not just an expert but it is a single agent so we consider that we have multiple agents like an agent that is reaching the goal an agent that it does obstacle avoidance

Uh in this particular Tod scenario but in more High dimensional scenarios you may even have like style based skills that you have learned from demonstrations caring actions and so on uh so by uh deploying a synchron like entropic regularized and balanced optimal transport uh algorithm what we

Can do uh we can compute the uh transport map of uh that you can see here according to a cost Matrix that relates to the problem that you try to solve which is like to reach uh a Target and adjust uh these VA parameters as adjusting the transportation map in an

Optimal transport problem and adjust what is the contribution of the different H expert agent and as you can see here we we have also some entropic terms so that we can ensure that we are equally sampling in the beginning all all the agents so that we can ensure

That we are trying to transport all of them but we also add the KL regularization also in this problem to relax uh uh uh the uh the optimization uh landscape and also ensure that we are not turning on and off uh experts uh abruptly what is interesting here with

This unbalanced optimal transport is that you can complete this switch of experts so whenever an expert is not needed you can actually completely turn it off of course the optimization decides on that and adjust the other expert appropriately and you can be extremely reactive and very very fast actually

This optimal transport methods uh solve optimization problems in milliseconds so you can actually handle uncertainty by being extremely fast in your decision making you can also uh uh deploy it on a 7 Dees of Freedom arm and even having uh moving obstacles and eventually you can also transfer it into the coordination

Of mobile manipulation is some we task that we wanted to to show here but we are very excited about this approach of uh uh multi-agent optimization for uh mobile manipulation robots or humanoid robots in general so I will move to the next topic that uh has to connect mobile manipulation to to

Perception on as the first step on how when we wanted to understand how we can represent Destin geometry in order to learn to interact with the world um so in general our robots have to act in environments that there is high clutter and classical uh basically uh approaches

In se3 grasping in the Clutter uh do not really handle partial visual information uh they usually have to scan the whole scene and then start quaring for 60 grass they usually do not use a single view unless it is just a top down camera that observes a a single tabletop

Environment and and does some sort of 4D usually grasping and um for sure many approaches are basically overfitting a single camera installation and they cannot uh be used at any random viewpoint but in uh mobile manipulation we need to be able to have partial to handle partial visual information we

Should handle single views so the robot sees the scene and already knows okay I have to grasp this object I have to go there and grasp it and not just start saring around and uh it should be able to handle any random uh Viewpoint and

Here you can see an example of a random views of of the robot H in in our lab with different installations of the SC so in our recent work is quite ongoing uh we reimagine grasping as uh rendering as neural rendering so uh what do we do

In this approach so we train a a neural network that consist on the first path by a occupancy a convolutional occupancy Network that takes a z the 3D tsdf uh grid of the of the single view of of the uh Scene It passes it h through um and

Extract a 3D feature through a 3D convolutional encoder and trains then an occupany Network that tries to uh Recon construct the uh the scene and this is trained easily by this um binary classification problem of 01 of whether a point belongs on an object in the

Scene or not so we try to do scene reconstruction and then we try to leverage this 3D feature volume in order to do grasping now once we have this occupancy Network we can go here and we can uh create um uh scene uh rendering of the scene deploying virual uh cameras

And doing gray marching and quering the occupancy Network to understand whether um what is the weather and point from the camera belongs to an object or not so in essence we can uh create a new Point Cloud uh compared to our initial uh partial uh Point Cloud that is more

Complete because we have done this SC Rec construction and now we can query uh grass we use a a geometric approach for quing graph based on point cloud and then for each grasp we go additionally and we do neural rendering as if the grasp itself was uh let’s say

A camera so we want to see what does the grasp see in the locality of the objects in order to extract features local features that represent the 6D representation of this uh graph and then we can train a grasic network that can uh classify grass CCH

Good or bad on a seene level so not on an level so if we go and look a bit more in details this is what I have already described how we train the occupany network it’s quite uh well known and it allows us to get the SC completion and

We can query also at any uh resolution our them so we may start with a very rough DDF and increase the resolution or vice versa but of course the question is how we can extract the right features for uh grasping so usually in grasping what is happening is that either you

Have a generative model that just tries to um by a variational decoder for example a query grass so it does basically the grass sampling approach and then you have to have an additional evaluator Network that evaluates whether or not this grasp is good or bad when you do uh grasp detection you assume

That you have some uh broad sampling uh strategy than like the one that I described the one that it is a geometric uh uh sampl uh strategy on point cloud and then you need a network that can evaluate if the crows are good or bad um however usually

This evaluation in the 6D space in sc3 is quite difficult so what is usually happening is that you try to evaluate the pose so the 3D vector and then you try to regress so you basically query tdfs uh the the voxels of the tsdf you try to understand if the voxel contains

A good grasp or not and then after doing this classification you try to regress in SO3 the 3D um rotation of the graph and this is also quite challenging because doing regression in s SO3 and also depending on the different representation is quite uh challenging uh what we propos instead is to extract

Features that represent directly the 60 uh pose and we do this by this uh rendering approach so we render the thing with sample graph on the on the new Point Cloud that we uh that we get we can arbitrarily create as much dense

So um SC as we want so we don’t need to rely on on the initial uh you know rough Point Cloud but we can make even more dense uh Point Cloud we can even H sample on this point new Point cloud with additional constraints of reability or semantics if we know about

Affordances and so on and then we select each grasp and in each grasp we place three virtual cameras on these three side of the of the grip and we uh do neur local neural rendering that we also supervise locally uh trying to extract neural sdfs so we try

To uh we can we try to additionally can you hear me sorry I I thought my connection might have been no we do we do hear you ah okay okay sorry um so we can additionally uh extract local features via this by this uh local uh surface uh rendering and

Extract this CRP relevant features that we can then use uh in order to uh learn uh the uh how good a grasp is or not uh additionally this local rendering comes with the benefit of getting a bit better also th reconstruction although this is not actually the the real um uh our real

Target in this work but rather to extract better features for representing uh graph and how this pipeline looks like uh so we observe we take a single uh view of the scene from a PR grph post this is uh the scene Point Cloud we render the scene then we do graph

Sampling then we uh pass the grass so we do this local rendering per grass we ex extract the features we pass them through the network that judges which are good uh grasps and then we execute grasps sequentially um and here you can see the simulated uh uh setup this uh simulation

Setup comes from an earlier work from uh 20120 if I’m not wrong uh from the volumetric graphing uh networks the was the one that was doing this uh classification of voxels that I was describing and here as you can see with the red is the rec construction yellow

Are the sampled grass and green are the good grass and compared to other approaches that are completely implicit in sixel or semi-implicit in the sense that they are implicit in evaluating the pose and they regress uh the the orientation you can see that we achieved uh state-of-the-art results in

This Benchmark also in other uh data sets and we can also transfer in the real world and here you can see a decluttering action of the robot here we don’t move the robot we just have this single view we we render and requery every time that we move an object and we

Basically obtain a new uh scene but we can also move our robot around change the cam Cera height because we have the Torso and uh additionally we can also select with which arm it makes sense to execute the grasp uh given that we have a mobile a by manual mobile uh manipulator uh

Robot um I don’t know if there is a question um so far otherwise I will move to the last part hear me yes yes I yes right is there any benefit to have a detailed read model like from nerves over occupy Fields is sorry can you can you say again the

Question if there is a paper there is any benefit to have a very detailed 3D model from nerves for example um so it is similar to to nerves what we do right but on the other hand we we assume that having the color uh rendering irrelevant to the

Grasping actually um if this answer your your question so in essence what we do here is uh Rel is similar to obser so basically you do just the surface uh rendering with without caring about uh colors um also because in my opinion if you have uh object uh with different

Color that can mess up basically your graphing pipeline because graphing is quite a geometric problem so you really care about surfaces and local geometries uh so in this case uh doing a true Nerf let’s say uh with the color uh rendering might mess up in some sense the features

Uh for example that you are trying to learn you add an additional uh complexity uh let’s say if we assume that knowing also the color in particular during the grasping generation I mean it is necessary and it’s not something that we can append you know by doing some RGB um I don’t

Know open vocabulary object detection cannot be handled elsewhere I think that um it is not really needed the color in this kind of pipelines I don’t think it’s really needed uh for grasping I is there any other way that you can think of like a way of incorporating semantics basically better than

Maybe so yes so the extension of this work that we are currently working uh considers parsing and trying to do semantic reconstruction of the scene and through the semantic reconstruction we are actually trying to predict additionally affordances for the grasp so not just uh grasping generation but additionally classification of the

Affordance of the grasping affordance so uh basically if you have a knife uh what is the usability of the knife so you need to do a Handover so if you need to do a Handover with the knife you want to grasp it from the uh um from The Cutting

Area and not from the grasp area so if you want to cut you cut you need to grasp it from the grasp area one issue that we have encountered uh so far is that they are not really good enough uh afford grasp aordance data set for uh

With 3D models there are uh uh affordance data set General affordance data set there are also in like shapenet a lot of these objects that they might have like random affordances like chairs that you can sit for example or like other stuff uh there are some data sets

For um articulated objects in which case you just care about predicting the articulation of the of the object but the general grasping uh data sets with different functionalities about the different objects for example uh battles with which you where should you grasp them if

You want to p uh or this kind of uh basically types of grasping uh affordances they are quite scarce so now my student has collected um a subset of um I think of 3D affordance net and uh shape net um and he created a an affordance data set

So that we can pass semantic information what we currently don’t see quite easy is to create something like um joint open vocabulary um uh segmentation Network plus of uh graphing what happens uh also now in in the robot Learning Community usually people uh just query uh separately with an open vocabulary model

And then you for example do grasping um on the highlighted let’s say uh object uh which is Trivial to do I will I will also show you like in the next work we also use like a a seman semantic detector decoupled from the from the actual execution thank

You thank you so the last part maybe I’m a bit late I will try to be fast um the last part has to do with active exploration but also H execution by mobile manipulation uh robots so we already uh discussed about mobile manipulators that they have to act in unstructured cluttered environments they

Have to use ideally only their embodied camera in the same way that humans use only uh their eyes I think that it will not be really nice to have robots in houses and have additionally multiple cameras right in our in our houses uh to inspect us so that they can help the

Robot so the robots would need to actively perceive and explore in order to obtain this task relevant information under reclusion and uh one question that we have in Mobile manipulation compared to just uh you know a navigation for the purpose of scene reconstruction is when do we have enough information so that we

Can stop the exploration and execute the manipulation uh task so mostly in active perception uh you will see uh mobile robots moving around just trying to do SC Rec construction and the goal is to represent as well as possible uh the scene and even for uh grasping with static

Manipulation the range of motion is quite limited the scenes are quite small while in Mobile manipulation the scenes are larger the motion that the robot might do for reconstructing the scene is uh more costly and we don’t want to do wasteful motions we want once you detect

The object of interest that you want to manipulate you want to execute the task so we need some sort of time and energy efficient solutions for mobile manipulation so we consider a setup where we have again our robot with its embod camera and um the robot doesn’t

Know anything about the scene the only information that it has it’s a rough uh quite large approximate Target area with this represented with this red bounding box it only knows that roughly there is an object that it has to find and some instruction like pick up the object at

The right corner of the table so the robot now has to explore the scene bu the tsdf of of the robot a volumetric representation and then do actic perception and Gras detection in order to understand that now I have found the object of interest and I can I can grasp

It um here we didn’t coule it with a uh uh SE reconstruction of the previous work ideally we would like to somehow bring these two works together and and somehow explore uncertainties from the from the previous pipeline inside active perception but for now assume that there

Is no really a network that can complete the scene so the agent has to uh really uh sample candidate paths towards the area of Interest this red bounding box again remember remember that we don’t know uh the rest of the scene so we broadly sample uh base poses and Camera

Poses towards this uh Target area uh the grass poses are sampled based on an approximate estimate of Target reability based on what I presented in the first part of the talk and we formalize this problem as a resending Horizon control problem or as like as a model predictive

Control problem where we try to execute the um best path uh trying to balance two utilities two objectives the one is the Information Gain objective uh and the other one is the executability of a possible uh grasp so we try to collect as much information as possible given

The the tsdf and the camera poses that are going to take in the next pose uh but we accumulate this cost this um utility along the path so it is not only next best view but it is a AC cross a discretization of the path till the

Final goal that we uh that we sampled and therefore we also found it very beneficial to actually penalize to the distance that the robot actually uh uh travels in order to accumulate this information and then on the other part we have the executability uh objective

That uh tries to um uh stamp graph uh initially of course there might be no grass because you have not found their target object but once you have the object become part of your site so if it is if you manage to go close to this um

Target area you can start query uh grasp and you can query their reability Additionally you can query their affordability for example if they really belong to the object that you want to uh manipulate because they may be grasps on a irrelevant object and again uh we try

To penalize the uh the actions that the robot will have to do in order to execute the grass so how much will the robot have to travel from its current location in order to execute the grass so we try to make everything uh quite efficient and the fact that we are doing

This in this resending Orizon or NPC style uh control allows you to com to reactively replan at every time step as long as you get like new 3D thin information in some sense we can see this a bit like modelbased um reinforcement learning where the a model is basically the scenery construction so

The robot tries to build up the uh seen information and then uh stop whenever it has found enough information and Gras in order to execute the grasp so here you can see an example of this uh approach so the robot samples poses uh for its

Body so the the motion of the base and in particular the motion of the head and the torso um we assume that we don’t move the arms because it is actually relevant as long as you don’t have to manipulate something it is just a waste

Of energy to uh try to uh solve I case already before being in the scene and also you don’t know exactly if there are collisions because you don’t know the thing uh so the agent uh samples and it can also do collision avoidance so we can add additional objectives of

Collision uh avoidance and here you can see the two uh utilities with light blue so with pink is the uh psdf that has been already discovered and with light blue is the tsdf that you expect to find uh if you move to a new uh here in this green uh basically uh uh

Head poses basically and this is the executability that contains uh possible grasp and uh the reachability not on the ground here we represent it on on the agent space for both arms in this particular case so here you can also select which arm to use the one that has

The highest ability in order to uh grasp the object compared to the first work that we were only considering the one arm and we have also seen that in particular if you consider affordances that are uh uh necessary to act on Sideways so to do side grph and not only

A top down grass have having this balance between Information Gain and equability kind of a exploration exploitation basically balance really affects the performance of the of the grasp of of the robot and ideally we would want to place the robot in a location that it can already execute the

Grasp and that the robot doesn’t really have to do additional um uh steps on deciding how to place now my body in order to gra so we try to balance everything in this uh single uh Pipeline and it is quite different to doing next best view that

It is very myopic again so it’s like a uh the robot jumps around in order to collect enough information as you can see goes randomly around the table without being oriented to the task that it has to solve compared to our approach and again it was almost easy to transfer

It into the real world and now I I I assumed I forgot to add an additional actually video here so you will see only one execution um so we did several attempts with our robot we have quite a very nice uh uh High success rate also

In in the real world however this comes with various challenges because as you move your robot and you want to collect information while your robot your robot uh moves the noise in the real world in the base affects a lot the uh the camera as you might EXP effect so here you see

Uh of course a times four I think execution of the robot uh but in several steps we had to basically uh discard uh some frames of the tsdf so we were not updating the tsdf at every frame but we were discuss discarding several of them

So as a takeaway uh message to not take much more of your time um uh what we have seen is that we can do efficient mobile manipulation if we just consider the co coordination of multiple embodiment and not necessarily try to solve the whole thing as one although

This is needed if we just want to do whole body manipulation task like opening a door and so on um neural geometric representations are necessary for skill learning and efficiently coupling the geometry of the body of the robot with the geometry of the scene is necessary uh and there is this crucial

Interplay of perception and action that we try to find the right tradeoff and I believe that once we find this connection between how to perceive in order to act and then act in order to perceive we will have the perfect balance to do way more challenging task

With more bi manipulation robots So yeah thank you for your attention and if you have U questions at this point I am happy to answer any questions I have some uh if it is okay like hi Georgia this is chai basten thank you this was a very nice

Talk I really enjoy it and it’s a very interesting work and I also like the name of your lab Pearl I think it’s a good name to remember it uh so my first question is related to the reinforcement learning that you presented so if I understand correctly you do that

Learning I mean RL agent learns to reachability uh based on reachability right like but like you may like so it’s like in terms of uh basically the position but how about the control of the forces which is like manipulability uh I mean you mention it slightly but I miss it like maybe you

Also consider it in the maps I so the the maps consider manipulability but we don’t consider force exerted forces in this work uh so this is like part of the uh basically manipulation uh policy in particular so uh as I was saying while I was presenting we are now extending this uh

Work into a perception based approach where we can uh effectively uh coordinate the manipulation of the arm while exer exerting forces for example when manipulating ulated objects for example this has to be added additionally into into the problem so here it is quite kinematic indeed still

Uh all the the tasks that I have shown here are quite kinematic and they are not contact reach where you have to additionally handle the excerted forces by the by the robot in the environment uh all right thank you I mean like I guess learning can be done by both considering reachability and

Manipulability together right yes yes you can add it as a reward function basically on your uh uh yeah you can add it in the reward as a reward uh as part of your reward function um in order to promote uh to go towards areas that increase uh manipulability of course

Again you have to scale uh the reward in the reward function so if the task for example is to open a door uh the manipulability has to be um much higher uh so that it can promote uh posing the the the body of the robot in a position

That will make it easier to open for example a door yeah okay uh second question is the second part of the work I work in the area of haptics uh you know human and machine haptics and I’m wondering like can you provide us some I

Mean if if you you’re not working in the are of Optics but like what is what what do you think could be the role of htics in this interplay between perception and action like in terms of grasping objects or like putting on a you know uh so um actually this is quite interesting

Because yes htics is a very important part for as a sensing modality for robotics I don’t believe we will most of the time is ignored I think that is it is ignored I think we don’t have a good enough sensors to really get the exact feeling like a human right so the

Classical Works would consider that that the for talk censor can only understand let’s say the accepted forces but this is enough if we as I was saying before like open close a door uh you can get like the normal and okay some tangential forces but then you don’t really know what is happening

Inside the the hand and it is quite interesting because yesterday I was uh teaching grasping to uh to in my lecture and I basically took them from the analytical model everybody was expecting that I’m going to talk for deep deep based approaches and just in the end I

Just told them some deep learning approaches so uh the overall discussion was that with deep based approaches like perception camera based uh let’s say approaches you can only H get quality of Gras or sample Gras you don’t know if the grasp can be maintained and you

Don’t know if the Gras is taable uh the moment that you have grasped the object there is no guarantee that this grasp is actually good in terms of force closure for example H and for that you really need to have feedback so the question is where where does this feedback come from

So we uh we really need to uh see whether or not with tactile based um sensors there is quite some work now in in some robotics labs for developing new types of tactile sensors that are quite cheap uh they are B Vision based also so they have like cameras uh uh in the

Fingertip and then there is like a gel soft material trying to resemble like the the human uh uh fingers uh but they are also challenging because for example you cannot do any type of grip you have to be very delicate you cannot exert any type of uh of force with your fingertip

Um and in general this is also for Ines is not for um um Power grasps for example for many tasks you have to do power Gras so you also this kind of things you cannot really do with two fingered uh grippers that we use so definitely hoptics is very important to

Be coupled with perception and actually I have a student working on multisensorial uh reinforcement robot learning as I call it uh for this kind of robot to to try to understand how we can couple more information but yeah if if you have any idea of any other way of extracting hoptic feedback from

Interaction even with the two fingered robots I will be glad to here okay we can talk maybe sometime later like yeah sure but thank you that’s very detailed thank you any other questions Georgia I have one about the last part you said that reconstruction is like model based right yeah more or

Less if you were learning the model right constructing the whole world is very difficult right like can you make it more construction do you have any ideas about that yes um this is quite interesting so um what we are trying to um to consider in uh ongoing work while

It is quite challenging so I don’t have any results to to to show you um even in a single room with a few Furniture it is quite challenging to to to create a representation of this whole scene so that the agent also can use it for manipulation uh so what I currently aim

For and I believe is something that is also inspired by humans so humans we create different types of map um so I think that we should have some sort of course and then find uh representation so the course representation could be either a very um rough voxelization of the word that

You can try just to to represent maybe as an as an occupancy like these feature volumes but it can also be just a topological map I am quite uh sure that having a high level representation as a topological map is very very important but then I believe that on the low level

We need h fine grained neural representations this geometric representations that like that these ones that I was describing so more local fine grade geometric representations that are relevant to the task that you have to do so I can imagine that if the robot has some sort of uh representation

Of the scene that should be possible semantic topological representation so I know that I am in the kitchen I I am now in front of the bridge of course with some uncertainty we have not talked about uncertainty at all and I think that an is super important also to uh to

Leverage um now I have to handle the fridge and then somehow you have to trigger these models that represent your interaction with the fridge and also this kind of uh represent local representation let’s say of the fridge in in your um in your view uh so you

Will we as humans always have these kind of mental models about objects uh but when we interact with a particular object we focus on the specific parts of the object so I think that this is very very important for uh more large scale mobile manipulation let’s say yeah sounds very

Exciting let’s see if it will work it has potential thank you very much Georgia Thank you thank you too well bye thanks everyone see you next week thank youor bye bye

Share.
Leave A Reply