
The Quantitative Biosciences Institute (QBI) at University of California at San Francisco and the Office for Science and Technology of the Embassy of France in the United States presented the “HealthAI Symposium,” a French American Innovation event at the University of California, San Francisco from December 6-7, 2023.
Supported by renowned research institutions, the health symposium brought together leaders, industry experts, and researchers from the United States and France to talk about AI applications for the health sciences.
MLtwist’s Audrey Smith joined Laurence Calzone (Research Engineer, Computational Systems Biology of Cancer, Institut Curie), Michael Keiser (Associate Professor, Institute for Neurodegenerative Diseases, University of California at San Francisco), and Hocine Lourdani (Group Product Manager, Deepcell) in the Artificial Intelligence for Cell Biology panel discussion, moderated by Moderated by Lisiena Hysenaj of Parvus Therapeutics.
Check out the full panel discussion (Day 1 of 2).
hello everyone welcome I’m jacen fabus the Chief Operating Officer for the quantitative biosciences Institute qbi
and I’d like to give you a warm welcome to today’s event uh today we’re kicking up our third joint event with the
scientific Department of the embassy of France and specifically the French
Consulate of San Francisco uh on the topic of research and health our first joint event was in
2020 it was virtual due to the pandemic and I’d like to say that we were slightly ahead of our time with the
topic being AI big data and health in 2021 it was still virtual um it was a
multi-day event titled reshaping science and health each year has been a great
success but I’m happy to welcome you to our first in-person event today in 2023
Health Ai and with this I’ll now pass the mic to our partner the French
Consulate of San Francisco represented by Emmanuel poak vour the atache for
Science and
technology good afternoon everyone um I’m Emanuel P I’m the ATT for Science
and Technology at the at the Consulate General of France in San Francisco it’s my great pleasure today uh to be your
co-host on France’s behalf I would like to thank our longtime partner qbi and
jacn for uh hosting us and co-organizing this event with us which is a French American event so you
will see a lot of French or French speaking people around and that’s normal I don’t think I need to convince
you that the topic of today’s Symposium is of utmost importance we’re proud that
France can contribute to these advances hand inand with a renowned partner like qbi I’ll try to keep the speech as short
as possible but I have a few things to say um one of our main missions at the French office for Science and Technology
of the embassy of France uh along with all the scientific attaches in Atlanta
Boston uh houon Chicago Los Angeles and Washington um our objective is to Foster
scientific collaboration between French and US researchers and support students
and research Mobility between our two countries especially at the doctoral level or early stage of the carea it’s
not just by chance that the collaboration between France and the US is is one of the strongest and the most
enduring our two countries do have a lot to share in terms of knowledge and knowhow in all aspects of society
including science and research we also value the supportive
entrepreneurial entrepreneurial spirit that exists in the tech world and which
is particularly strong in the Silicon valet as some of you may know so incubators startups spin-off companies
and large Industries are the ones that truly turn the most mindblowing Research
into a reality that will impact and improve the lives of the greatest number we could not be more thankful to
see French and American actors of innovation come together in an event like this one and none of these would be
achiev none of this would be achievable without their kind participation so I really want to thank them
truthfully my final message for today is that this Symposium is a good example of
what what we can achieve with all of your support uh as well with the French and
the Fran ofile community so it’s only one example among many initiatives
supported by the French Consulate on an everyday basis our goal our job is to listen to to talk about
the scientific and technological subjects which are of crucial importance to you and to both our countries and the
future of Healthcare in the AI era is definitely one of them we aim to create
dedicated spaces in communities where scientists can meet share major scientific results and insights that
will impact the research the economy and the society significantly there are many ways in
which you can support our actions whether you are French or not so please
come and talk to us during these two days learn more about the things that we do and the many ways that you can get
involved and just a quick word to maybe mention the Photo Expo that you you may
have seen when you came in this room which was one of the initiatives of the office for Science and Technology last
summer and we actually had some laurates uh who won some prizes in this photo competition so next year please
participate without further Ado I leave the floor to Tanya CM who will also
introduce our first panel discussion and I wish you all a very fruitful and enjoyable Symposium thank you
yeah good afternoon um and welcome from me um as well I am Tanya corta I am a
professor here at UCSF in bi engineering and therapeutic Sciences I’m also the vice dean of research for the School of
Pharmacy and personally extremely excited about the Confluence of AI and
Discovery Science reaching all the way from really fundamental curiosity to to
impacts on our everyday life and so to me some of the most exciting advances
happen at the intersection of Science and Technology of the intersection between different disciplines in science
and at the intersection um between technology science and healthc care and
so I think that’s why to me um what we’re witnessing um all over the world
but also what we’ve been discussing in the last day and a half uh um in a sister Symposium on AI across biological
scales and what we are continuing to discuss today in the panel discussions um will be really at this exciting
Confluence um of advances in computer science in cell
biology in drug Discovery in health care and in clinical applications and so I’m
very excited to hear from experts in this area um how they are viewing all of
these um developments and what they can inform us on what the future holds in
these areas and so with this I’m going to introduce um our first um panel
moderator um lisana heai to introduce her panel on
um cell biology and Ai and also the panelists welcome thank you very much
so the first panel will be AI biology and we’ll have four
speakers the first one it’s Michael Kaiser from UCSF he’s an assistant professor at UCSF I think this
microphone is not working okay yeah
okay and the second speaker is uh Laurence from um Kon from she’s coming
from France from Institute C and then we have oin Lani that is an engineer on AI
and O Smith from ml twist so every one of the speakers will present itself uh
for three minutes and then we’ll have a panel discussion together so maybe Michael you can
start
all right maybe we’ll put this all right well thank you so I’m an
associate professor here and my background is originally in computer science and over years I actually came
into this area through uh through UCSF itself I was once upon a time a bioinformatics graduate student here and
so the lab and we’re giving very brief ENT um our lab focuses um on three major
application domains in bio medicine and biological research for us Ai and ml
questions are at the center of it and so we look at molecules and we spoke a little bit about that yesterday about
perturbations and here when I think of a perturbation it could be a molecule it could be an environmental Factor it
could be a functional genomic change to a biological system and so some of the areas we’ve been interested in have been
cell biology as relates to this panel also whole organisms and collaborations
in zebra fish and looking at neuroactive activity and then Imaging Imaging can
span so many things in our area of Interest here I’m showing some examples of collaborations in the space of
digital pathology where we work on becoming a force multiplier or a way to help Pathologists operate at a scale
resolution and generalizability that otherwise would be impossible to do as a
human and so we were asked to show one or two examples I’m going to take the easy route and show two that were from
the presentation I gave yesterday the first one um is a generative AI model
where we were ingesting images these are microscopy images from cellular screens
for a neurodegenerative disease phenotype specifically the tow protein
and at the left what you have are two fi are two different channels of the same field of view of these cells
at the top you see yfp fused tow protein and at the bottom we have some nuclei
now as I mentioned at the time one of the big problems with live cell Imaging to do drug Discovery and other screens
at scale of hundreds of thousands or more perturbations is we can’t put antibodies in that system we can’t run
it in a fix cell format it doesn’t scale and so if you’re doing this in live cell and we also look at different time
points we’re stuck with only seeing all of the tow what we really want to see though is the misfolded tow or the
Tangled towel this is the tow that builds up pathogenetically and we see it in the brains of patients who have
passed away from Alzheimer’s disease and that’s only some of it and so what we were able to do was train a deep
learning method that could take data from screens that had been done years before and ingesting those two input
images on the left generate what we wished we would have had all those years ago and specifically focus on the tanks
Town happy to talk about that and then the other sort of result that
illustrates area of interest for us is a second is a second generative approach and so this is a generative diffusion
model and what we’re able to do is generate fragment by fragment piece by
piece molecules within protein pockets and so combining these two approaches
you can imagine a whole pipeline of testing experimentation and assessment
in order to move forward understanding cellular diseases and Drug Discovery
thank you thank you very much and so now we are still in Academia and we’ll have
someone that comes from France L will introduce what she’s doing and who is
she thank you so my slide are not as beautiful as as what I hoped but um so
my name is Laurence Calzone I’m a researcher in the Currie Institute um I
coordinate together with Emanuel bario with in the audience the group of computational systems biology of
cancer um I’m a mathematical modeler and with the focus of uh on cancer
applications and what do I do what am I interested in is to build mechanistic model that describe the biology of the
disease and in our case in particular cancer but not only with the purpose to
understand the impact of of the alteration suggest optimal drugs and predict drugs response or drug synergies
in a patient specific manner um so where do I come from so I’m
a mathematician by training uh I uh I I started doing uh chemical kinetics in
particular and then I move to biology through Theory and I did a PhD in theoretical biology here in Virginia
Tech not here but in in the US uh so my first interest were to do some dynamical
modeling uh of uh of the cellular processes and uh and then when I moved to K uh the
the type of data we were dealing with could not be applied to ordinary differential equations really so this is
when we developed some some other formalism uh based on on both um
approaches but what we we really develop in the group are models of tumor Evolution and expansion uh taking into
account space and time so with have a a big big data and and a lot of parameters that we don’t have uh we also developed
some tools to model individual cell populations and we want to integrate of course all the omx data that we can into
our models so um I’m really interested in
mechanistic models and how to make the link with artificial intelligence models you know the we try to uh explain really
the type of results we get with AI models and we can do that from two perspectives first as an input to the
models so we learn from the data to build our mechanistic models of course with the stratification of patients but
also help us infer the models that we build both the networks and the the mathematical equations from this data we
want to integrate multimodel uh data as model parameters so again learn the parameters from this U from this data
and of extract features from the data which PA Pathways or which genes really separate the response to a to a
treatment or not and then when we have this model we are also interested in in in uh um AI to um improve the models and
how do we do it uh we search for optimal parameter sets so we have a lot of parameters that we need to infer and
this is this is really difficult to do it by hand and of course we expect to to use um uh AI models like surrogate
models to mimic uh the complex and Compu
computational heavy models that we build and we also use AI to generate synthetic
data for model optimization we don’t have sometimes we don’t have enough data and this is one way to augment the data
so we want to make the link between AI results and model outputs and really fill the gap between mechanistic models
and Ai and I think I will try to give this perspective in uh in the response to uh to the the questions that we will
we will talk about um the point is to explain the content of content of the black box as much as
possible and and I was also asked to give two uh two um um projects on which
we work in the lab the first one is uh is called is um the point is you know it’s the holy Grail to try to find
some biomarkers for a response prediction to to some treatments and here it’s about immunotherapy and we use
multiple data types to uh increase the the power to be predictive and explain
as much as possible uh why these patients respond or not and this of course open the road to special modeling
which brings me to the second project uh which is my expertise is spatial
modeling of tumor Invasion we model both intra and extracellular interactions um to uh to simulate in
silico treatments so we start from data we infer the model we infer the parameters and you can see here an
example of of a tumor and and invasive um that starts invading the the tissue
and then we can play with these models and search for optimal drugs in silico that we can then test um in uh in
experiments and I will finish here okay thank you now
it’s oin it’s a farm D that is going to explain us more what he has worked for
some companies in the Silicon Valley on AI so he will introduce
himself sure hi everyone thank you every thank you for inviting me and thank you
for being in the room today really appreciate it quick int introduction by myself my name is Jos Lord I am a farm D
indeed By by Train um started I am originally from France which I wouldn’t mention uh usually but feels relevant
today i started my career in in France actually in Pharma in biotech then went
to business school moved to the Bay Area a bit more than 10 years ago and since then mostly spent my time building AI
products and so I I’ll try to bring that perspective of a practitioner today and um and somebody who has led the
development and the commercialization of AI products in different Industries in
part Healthcare and uh and Life Sciences Industries so a couple of examples uh I
I extracted here uh is the work of uh I’ve conducted with the teams at arteries
and Rapid AI where the goal was really to leverage AI in a clinical setting um
to build software medical devices to support uh clinicians in their diagnosis
and treatment decisions so very specific set of challenges are related to that
it’s really about the measurement of certain clinical uh factors the
detection of certain pathologies and that’s a very specific type of of challenges that I’m sure will come up in
our discussion most recently I I’ve worked at Deep Cell where we took a very
different approach and Deep Cell is a company focused on Research especially at the Single Cell single cell research
and we applied AI to study the morphology of cells so we built an in end platform um able to take images of
cells in flow and extract High dimensional information from these images and help researchers characterize
and make sorting decisions U using um using the outputs of our foundation
model so very different set of of of challenges here and so I extracted just a few thoughts that may color my
intervention in this panel very high level and I won’t take too much time one is um I’m a product person
so I’m really focusing on building uh products and solving concrete problems
as we know here these days there’s a lot of hype a lot of excitement and for good reason around AI um but beyond buzzword
beyond the hype um I tend to focus on identifying very specific concrete problems and focus on the impact that AI
can have on solving that problem uh with an emphasis on the inter interpretability and the usability of
these models because the best model in the world can be very is not very useful
if um the the end user cannot really use it in good conditions um the second
element I wanted I kind of touched on is that challenges in AI develop product development and commercialization
implementation are very different depending on the on the industry uh in particular as I mentioned clinical uh
environments versus researchers only research focused environments are present very different types of
challenges and that may uh color my intervention today and finally one point that I wanted to to insist on and
actually that resonates a little bit with what Lance I believe was was presenting is that one thing I’ve observed in um recently in this and in
my my time working with biologists and researchers is that there is a persistent cultural difference Ando
difference in approach between researchers in biology and Ai and the AI community at large and it focuses a lot
thought and my experience in explainability and in whether or not a model actually describes a mechanism and
of course this is a complex topic so I W this is way too high level but I want to emphasize here that sometimes that qu
that question is very important and of course a model should be explainable and at the same time the value of a model
can be in its predictive power alone that happens and an example that I didn’t work on personally but I think
many of you in the room can relate to or or know of is Alpha fold for instance this model trained by Google Deep Mind
to predict the 3D folding structure of a protein based on its sequence of amino acids there is work on going on
explaining exactly how the model works and how it’s making its predictions but the value of the model in unlocking
avenues for research um just by being that potent and that that efficient in
predicting the structure is is already huge and sometimes explainability is not um the the only
aspect that we should focus on so I’ll stop here I took already too much time thank you everyone and I look forward to the to the
discussion thank you and so now we continue with other four panel and she’s
alre she’s a woman founder and she has build this company that is on AI
Services she will explain more about it hi everyone thank you for having me
so um two facts about me I’m French as you can probably be here uh and the
second one I’m the only non-scientist person in this panel so very uh very happy to be here very excited to be here
but I might offer another uh another point of view um so um I founded ml twist uh with my
co-founder three years ago and the idea is really uh to help soft a very um very
boring problem actually uh but that is like very persistent when you uh start training a model is um on uh the fact
that there are still like manual task that still need to happen when you need to create data processing or data
augmentation and when you try to not only build data pipelines for certain model but also maintain it because it
can break over and over again and that type of work actually um is taking 80%
of the time of any data scientist which is which is a lot and most of the time if you talk with data scientists what
they want to do is spend that time training a model not doing that genitori work that we’re happy to take
over um if you look about like the yellow shun you’re going to see that as
of right now a lot of um Technologies happen a lot of technologies have like raised over the past decade uh but what
really is not really like you cannot see happening is that you need to build the
data pipeline between all those yellow box to make um to train your model and that’s something that still not have
been like really solved as of today and that’s what we focus on if you look at the work uh the manual
work if you have a technical person on hand in your company or in your uh in your work um in in Academia you’re going
to have to go through all those different steps uh to get your data ready for a machine learning model to um
ingest the data and get trained on the data that can take up uh like several
weeks as you can see um and we build a platform that can auto generate data
pipeline um and all the uh purple uh steps that you can see are done
automatically by our platform which um um help people like you or data
scientist to just focus on you know like the exciting stuff which is uh working
on the accuracy of the data or just training your model and and checking the performance of the model
um as you can see we have been uh lucky enough to partner and uh have also great
customers that we allowed to talk about so we are working with s National Laboratory but also UC Davis health um
we are also working in uh with Stanford uh HAI department and um partnering with
Cap Gemini and and other uh names on on this slide um and so we have been able
to work work on a lot of different data types uh from what I was talking about
like images but also videos um it can be also text and it can be also dicom uh we
have actually worked on daos which is a spin-off of dicom um and that’s about it
that’s us in a nutshell oh you can you can see like a bit of our accomplishments too um
but yeah that’s that’s about it thank you thank you
so okay thank you very much everyone so as we saw in this panel we have uh
people from Academia that work really on building this model and making sure that we’ll have some discoveries from these
models and then people from industry that are making sure to create a product that can be used for everyone so now the
question I have it’s what is one of the
most impressive discoveries that we have had in the field of AI and cell biology the
recent years that we could not have had just with humans alone anyone wants to answer this
question maybe can you hear me yes maybe I’ll get started just because I just
reuse the example I took actually I mean in my presentation um I think Alpha fold
is is one of the most impressive developments that we’ve seen lately um
uh the it’s interesting to see that it’s a collaboration with Academia initially with emble um in Europe to build the
database and then the development of the model um you know was conducted by
Google deep mine U to your question Lana so the the problem of protein folding
and the 3D structure the prediction of a 3D structure of a protein is a well-known problem it’s been worked on for decades it’s been worked on
tirelessly by a number of groups and uh I’m sure many of you are in the room are aware of it so I won’t spend too much
time of it but the on on it but the level of performance that was reached recently in 2018 with the first version
and then 2020 um is just something that we couldn’t have um reached without the
latest development of AI um now you know I when we say would would not be
achievable by humans I’m always a little bit you know not sure what we mean we always use mathematical modeling we
always use different tools and these are created by humans and AI are also created by humans so I would still you
know put that on on the credit of humanity uh but this is I think a very
solid example thank you so another thing this is like maybe in the mind of everyone it’s we have seen all the
Sci-Fi movies where we see that okay something it’s taken by a malicious
attack and some criminals took over our mind that in this case it will be our
machine learning and I was thinking this when I heard the talk of Michael where you were talking about deep learning uh
techniques that you’re using in the lab and how you can determine it’s like a towel Protein that’s misfolded or all
this so how you make sure that someone will haug your like how how are we sure
that what do you do to protect your system and someone imagine someone comes
Haws your system and then your students or PhD students or post dog will use a
model that is not right and will claim discoveries that are
false ah well so this is it’s an interesting question too because in AI um there’s so
much of a history of releasing our code this is crucial if you want to publish
at a computer science conference um if you want to do any of these things you have to make what you’re doing available
and we’ve seen questions about this with things like chat PT and the question of guard rails or improper use but when we
think about drug Discovery um there have been papers where people have asked what about old techniques what about things
we’ve known for decades computationally how hard is it to misuse a model and that has always been
possible it hasn’t been the computation that’s kept the the computational barrier usually that’s kept someone for
making a very bad drug it’s been the synthesis the construction all of the other steps where the rubber meets the
road where there’s actually a lot of oversight and a lot of institutional support and that is one way that the
models themselves are less at the center of that question um but what you are
afraid on it’s not the the model it’s like this malicious attack don’t exist it just like we are scared of it I am
less afraid of models than I would be about how people choose to use things okay and so for me I’d always take that
answer to our organization do the universities the companies oversight
accountability that would be the place that I would ask because I can because the flip side of this is I remember
years and years and years ago I was um a student and someone was presenting on
some new technology or result a professor who was visiting and it wasn’t
published yet they were hoping to submit it to a really high-end journal and someone said but aren’t you afraid that
you’re by telling us all about how to do do this you know someone’s going to scoop you or use it in some way you don’t like and he said young man um I
would be thrilled if someone would try to steal my work okay great so thank you
for your answer yeah no just one thing about I mean in Academia we share as much as possible so we we put on GitHub
everything we can whenever we can to make sure that also we get the paternity of of what we develop but also uh like
you said we we hope that people can use it and then make it better and then we all work for the same purpose okay thank
you so one question I’m having now because I was searching a little bit
about ml twist and your collaborations and so that I thought ml twist is a
service and it’s used only by companies but I saw in your website that you collaborate with Stanford and Berkeley
and I was pretty surprised about it can you explain how you help this big
academic institutions that generate a lot of discoveries in AI to improve
their data sure so um we have like several ways to work
with them or to partner with them uh one of them is actually using uh the
students at those universities that like our medical experts for instance and who
can help us uh do some labeling uh for some some models uh we have other things
like you said like for Berkeley labs it was a different because we were working on um uh a grant uh for the Department
of energy and we were able to collaborate with them on a plan to win
actually the phase one of those uh of the that beer um and every every time so
we also working with 10 for hii and that’s something else like they are working on uh creating some models some
llms uh internally and we are generating the data for them so um that’s that’s
like we are very very proud of that uh but like every every time we talk with an Academia it’s like going to be a
different need uh but we’re happy to explore any way we can partner with uh Academia okay thank you so basically we
we saw that Michael said that everything the discoveries are done and the models are not misus as people use them
properly and then we have some people to help label well the products in order to
generate even more discoveries now as a biologist or immunologist that I am the
question I am asking it’s I’ve heard all this different models and companies that
exist is there any way I can learn because I don’t know how to code I can learn how to use or where I can be
informed about this new technologies that are coming
out um it really
so I guess I’m coming at this question from a again product perspective I tend
to believe that not everybody needs to be trained on the technical aspects of AI and that is the role of companies
largely but also of course the contribution of the academic Community to come up with tools that are very easy
and intuitive to interact with that in the background do very complex things and use a lot of code and use a lot of
very complex mathematical concept cep but from a user perspective just you know kind of Shield all of that from
users um I don’t see a very I may being correct and you may have a different
opinion I don’t see a very strong argument for AI to be any different than software engineering in general for
instance we’re not asking everybody to be able to code as you just said uh in research and I don’t see why that would
be different from AI um I think the argument is a little bit different when you think about students however uh when
we think about students in the fields of of biology I think the the the intersection of computational biology Ai
and my biology uh in general is reaching such a level that there is probably a
better argument there to say that the Next Generation in general should touch at least to some degree uh AI training
uh but it’s a field that’s evolving so fast anyway that you know it’s something that is a very complex convers you know topic and I don’t have the the expertise
on it thank you no I would say that uh the key for any project is collaboration so as a
cell biologist you don’t need to know AI but you need to be surrounded by people that can help you and I talk the same
language as you yeah so this is the question it’s like where do I find because every time I go Ai and then you
have this courses I’m like oh I’m lost what I do so the question was like where can I find this Char GDP for example was
all over and everyone was talking about it so I started to use it and it was easy and save my life but where is there
any website or something you well I I would emphasize that the collaborate point because that’s what I was thinking too because actually I don’t think chbt
is the right metaphor for how we often want to use AI for high stakes problems
because instead I I think of you know someone who uses the centrifuge without
training or the fact machine early on and now it’s gummed up either you’re
working with someone to at least get yourself started or there’s a core facility or there’s a way to get the
training okay and the problem with AI is let frequently that we would gum up a
machine or do something to the rotor it’s more commonly that it would appear to work in a way that was misleading or
not useful because it is a pattern recognition thing it will always find patterns and exploit them so collaborate
collaborate thank you and so uh I’m wondering just like you are all open to
have a contact from someone from industry and build a new model or help them on a this is this
is great I think it’s a message we have to the show so I think we are I don’t
know about the questions or I have five moment okay perfect so you’ll tell me
whenever yeah I have more questions so another question that I had it’s okay
we’re in the same space we need to be trained do you think like a training of AI like how to use Ai and not to misuse
it should be mandatory in every company now or in every lab like the safety
training um I’m not sure exactly how it would look like but I think that it
would be very important to start educating people on the good and the bad
um starting by data bias uh which is like something that is like as human we
we all have biases but if we are of it if you let people know that they need to
pay attention to the data they’re going to be selecting for their model um then there might be a chance that we’re going
to reduce the amount of bias happening in the model um and then to go back to your question where you were talking
about like learning how to code on on what it means to do Ai and so on I think
it’s actually important to keep that separate just because you’re going to have people from different backgrounds
collaborating and they will fight also biases that way if you have everyone you
know educating the same way learning the same way in the end you’re going to have a lot of biases okay thank
you just again just to say that I mean in the company you need a wide range of
expertise again so you need to have people dedicated to that and the other ones knowing the limitations of of the
results and how to use these results is really key I think this if we have to all of us if we have to learn something
is the liit ations and the biases to be clear about the biases
yeah and just one thing to say about it um because of the strength of all the
different expertises and opinions we don’t all need to be programmers or developers of AI what will happen more
and more is these tools in different ways make their ways into the workplace into the research environment all of
these and I worry that it seems like when you read articles when you think
about and you talk to people who have questions outside of this space it seems as though it’s categorical do I trust AI
or do I not trust AI um do I use a prediction or not but when you think
about things like chat GPT which I I’d like to you bringing up earlier we are used to speaking and writing and we can
critique it we can look at it we know sometimes when it’s wrong and if you try to look for that you find it more and
more and so I would say the training that would make the most sense is how to be a critical user of it yeah and what
techniques to have to test whether this thing is telling me what it so confidently is
presenting I completely agree with that point and to take it one step further I
think one one thing that we need to put on the agenda even more is the question
of Open Source uh in the industry in particular because indeed in Academia and we mentioned it earlier there is
this culture and this requirement to share your your results to share your
code to share your data and in the industry it’s not always the case and for good reasons of course if you
invested a lot of money in Training Systems you want to protect your IP you want to protect your first mover
advantage in the market and completely understand that but we need to find ways to still have enough data available and
enough source code available for researchers in Academia or or not to
actually validate these models and and um and test them and be critical about
their performance thank you so basically I’m more optimistic now on the use of
the AI and the models so now just to talk about cell biology in general what
is as as an immunologist I’m more like in characterizing the cell types and
it’s basically what you have done a little bit in Deep Cell so like will we be a moment we will just recognize cells
by uh only one feature just for example
Imaging or we’ll show could still do our flowetry with 300 different colors and
then analyze by ourselves and how how far is this I’ve heard also like in in
the same space I’ve heard also about wet lab so if you can explain this and where
we are in cell biology right now and what the Yeah I can draw from my recent
experience at Deep Cell to at least answer the first part of your question um so we are at the point point right
now with the the performance of AI models on Cell Imaging and we can really extract High dimensional information in
a way that is consistent an improper way to to say that would be to sequence the morphology of cells and at at scale
right and that’s completely new and that really challenges or or subverts let’s
say the the existing methods in facts and fetometry um where I think uh it’s very
interesting is that currently the way to characterize cell is based on the taxonomy in biology that is in evolution
but still very much rooted in surface markers cell markers in general and
getting out of that tonomy initially is a little bit different when you start looking at phenotype and that’s the the
next Frontier is really looking at phenotype for itself and really leverage that as a new omic basically uh type of
of data um and the next frti after that I think is multimodel you know research where you integrate that data with
multiomic information from different sources um so that’s that’s I think how you you get to a holistic view of a cell
that takes very different um you know lenses to analyze it did I did I answer your question I I think so maybe Michael
has something to add about I well I very much agree with what you’re saying and what I was also kind of reflecting on
and thinking about this is you ask kind of when we’re there um and I remember um
1965 a paper on strong inference quote from it is the measure of a method is its use and so we’re talking about
methods and technology and platforms AI or otherwise or combined in this hybrid system for what use and so for some uses
we might already be there right now and for some things we we cannot see them by
Imaging without the right marker and we don’t know that yet and so there’s this interest of multimodality how many
things we can bring together without breaking the bank on the experiment in the first place yeah and so I would
actually turn the question back to you and ask for what use um and what is sufficient and I I was thinking more
than Discovery and cell biology for example could be used in the biomicro space so for example I would imagine an
AI um like model that could help me
realize what cell types I have based on some phenotypes that I see on the
microscope or what can be the molecular pathway that
are affected on this phenotypes and I saw that you are doing something like
this maybe in the lab to link this maybe also in Pharmacology and this is what
how you see this feel and what are the discoveries in
there so this could be an hour conversation I’m sure others you know have thoughts
on to in there’s big differences between supervised and
unsupervised we’ll go and see what’s there based on the of information we are
already toed up to collect say images ands compared to I have a really
specific need a priority bringing in my human expertise I know that none of those image channels are are likely to
get the job done and if they did I’m actually kind of that would actually be a good counteract test to make sure I
couldn’t get that out that maybe I
need so would best way can
is that’s what we as scientists
allim as as
possible any have question I think when
I R te so how how this
works
resources question canor from cyber and so in cyber you never know
your computer all the best
practices down ACC don’t want never and
that’s what about else C
guelin scient
unless unless you a lot of to unless you scr dat and the model
itap
tocce controls now we needer controls because we will try to do what we ask
and that’s the difference and so that’s what in Cy
sec
um so not that M because we are generating the data we’re not generating the models but that’s definitely
something that needs to happen uh for our customers so we work with defense
for instance and that’s definitely um mandatory to to have like
some ways of trying to break the model um but you can do it yourself you can just go on chat GPT and and then ask
questions again and again and you’re going to see at one point it’s going to break and you’re going to even see some
biases appearing like very quickly on what the pilot is it is a man or is it a
woman um or nurse um and and you’re going to see that this is actually all
coming from how the model was trained and what data was was you know used to
generate the CH GPT so that’s very important and that goes back to what you
were saying Michael that we need to be able to understand that whatever we see
that is AI stamped we need to criticize it we need to be able to look at it with
an eye of hey is it real or is it bias or can I really trust what what it does
or what it says um and that’s like something that should be done even uh at the school level like with our young
kids uh because with like everything that is gener it right now we don’t know
what is true and we don’t know what is uh not true and that’s very important to train kids and students to keep that eye
uh out for for that
do no I agree but this is a scientific reasoning right you you want to see how far you can go and and find the
weaknesses of of your models and this is something we do uh even when we read an article we don’t we don’t take it as as
it is we want to question it and I guess this is what we need to keep in mind with AI models they’re not the truth
they they get closer to what we want and they answer a very specific question not Universal questions thank
so now as as we heard we not be scared of
thei they work in the same way as we work for experiments so they check every
step and and also they’re all working in a collaborative way so we should be afraid
to go and ask them if they
industry and
then and everyone if you have any questions for theel could should line
up there will be a microphone here so you can line up and ask
questions I start thank you so much for the panel uh so I have a question
because I see a lot of like academic or pure Tech kind of players how do you
work with the Biotech Industry how are your Solutions implemented do you feel
that there’s any bottleneck um from the pharmacetical industry to like adopt all these great
Technologies right well so one of the best answers I can give is the second result I showed with the small molecules
is actually a joint development with Genentech and so jentech um Partners came to us and we worked on a close
collaboration for more than a year here where there’s both Mutual support the IP
is worked out ahead of time by the university and Genentech and by setting that road up nicely then what we’re able
to do is get real feedback because one of the things that stood out to me in that collaboration was we made the
molecules they looked pretty good but our our in the early version and our collaborators said yeah but that looks
really strained like super strained in practice we’ve tried a whole bunch of public techniques and and the molecule
it has an that’s just completely UN unworkable in practice and so it’s
having at least a champion on each side and people working together directly
would be my answer to you on that no I can just say that we are in
France we are promoting a lot this uh this PhD program where you have industry and uh and Academia working together so
we are really uh try to uh teach students to embrace space industry and
uh and think about all the the positive things you can you can have like you know access to clusters access to money
to experiments that we may not have in
Academia one more thing I would say is that a lot of the startups that are trying to bring innovation in the field
of AI are often spin-offs actually of academic institutions and so that in Industry uh Academia collaboration is
almost you know inherent to a lot of the Innovation that’s ongoing now your question was more towards Pharma and
biotech at a large scale I think the Pharma and biote can be a difficult
industry to penetrate initially when you’re startup they’re looking for usually a lot of validation a lot of
very early data however I think in the field of drug Discovery in particular AI applied to drug Discovery you see a lot
of different Partnerships that are actually ongoing um Asen with benevolent AI for instance is an example okin with
sop sop is and another one there a number of of of Partnerships of the of that nature so I think we’re going in
that direction more and more and fora companies realize that they can embed uh AI um models uh and products very early
in their development so you you think that in five years from now there will be more and more drugs that are in phase
two phase three that have been like discovered thanks to AI definitely yeah
this is already happening they’re they’re already in the pipeline but very early so we’ll have to time will tell
thank you thank you I might because uh just a
quick comment before there is a lot of ethical consideration and I will be hosting the ethics and AI next week and
I see you know like it’s very tomorrow and I it’s very relevant but anyway my question was more as a sale biologist
and we we we generate more and more like big data set like omx data and and all of that and we build model and test
those models um but kind of on the other flip side where you were mentioning like
usually we want like an end product and the end prediction with AI but can we
flip things around the bit and use AI to help identify Gap in our knowledge and
how should we build framework to help actually scientists trying to make sense of those very large data set that are
very complicated and have a lot of connection and help hypothesis
generation if you have any opinion on that
all right I think there’s two levels to answer first answer the kind of glib one but one that we want to get to working
is all of the explainable AI techniques interpretable methods saliency mapping
oclusion or even building models that operate through rationale layers where you’re having it solve multiple tasks as
a series of steps that are fully differentiable all the way through and if the model works or does not work
through that that information gate then that tells you something because you can change what information the task has to
pass through and so that’s one category of answer a a technological answer an
architecture style answer in practice much of the saleny mapping methods
integrated gradients guided grab cam either bring in their own artifacts or smear out signal and can only tell you
really basic things like is it looking at the dog or the background but trying
to understand about the ear the know is iffy at that resolution and in molecules goodness gracious me we don’t even have
as much intuition and so that answer is where the field is trying to go methods
wise the flip side answer is is a is a simple one and it’s a logical answer instead which is this is just one
technique and if anything use very simple models wherever you can standed
up swap in and out statistical models and other models and ask whether the bang for the buck is worth it to bring
in something that we have less of an understanding of and so that goes back to being scientists and the fact that
all of us here in the room have some have agency in making that type of decision without having to rely on on
hoping that a model that learns for reasons we don’t know can tell us something new that we didn’t know to look for so it’s considering it as parts
and swapping the parts no I I I would just say that maybe
uh one way to look at it is to use these models to validate an intuition that you may have so the question will be clear
it will not identify gaps but it will confirm or infirm uh an an intuition
that you may have one last thing we mentioned earlier
in the discussion the difference between supervised approaches and unsupervised approaches I think that plays a role in
your question when it comes to identifying gaps building taking an unsupervised approach to look for high
dimensional uh structure uh global and local structure in high dimensional space
using what is in many cases now known as Foundation models so models that are looking very
generally at at a field and then can be trained more specifically for a given application is an is an approach that’s
very relevant you can look you can really that that’s what I would say just from the the the ml training approach
just being unsupervised or self-supervised yeah helps a lot thank you I like that you said
intuition that likely scientist are not going to be replaced
by hi so my little thesis was in proteomic and omic data at Stockholm
University I wonder One Thing how can you deal with the biophysical Pro
properties of a shape shifter proteins like Spike because we had a lot of
problem to read the cre data talking of a particular class of proteins that they
are able to change confirmation because of
pH the short answer is no one’s dealing with that yet the ways you could go about trying
to start would be to think about your representation and because when we start
a model that in order to do any deep learning or machine learning training we
have to make a decision first and this gets tol twist for instance and thinking about what our representation of our
information is not just the foul format but what is this this cloud of atoms
what is this as far as the model is going to ingest it by the time it sees it for the first time in the world and
so geometric neural networks might be good you can make graphs spatial graphs
in space and then who’s to say that a graph only has to operate in space it
can also go across time you can put edges together to have on embl so that’s something similar to like convolutional
approach but in multiple dimensions and now you’re convolving in non-cartesian space so that would be one way to do it
you can use attention as well um so there’s a lot of methodological answers oh but yeah I want just to tell you that
for desperation we started to look at possible binding sides for transcription
factors oh yeah well let’s continue this after because there’s a bunch of neat
methods that can be done with this thank you
good afternoon uh wonderful discussion for the planel my name is bushan I am a bi engineer and I develop bi materials
and cell therapy uh I’m new to um AI in B engineering so pardon my knife
question so I have two part question one is uh what’s the panel’s take on model
hallucination and how it impacts uh uh the the uh clinical data or clinical
output uh or in in output in cell biology as well and the second part of question second part of my question is
uh we are talking more about going high higher dimensional in terms of analyzing object or making uh new objects by the
by the means of generative AI uh but in terms of clinical translation since I work on translational side of a science
uh we want to have a simplistic system to be uh clinically translatable so what is the panel take on dimensionality
reduction uh going forward after we have a lot of higher dimensional data
in based on the nodding um the first part of your question sorry uh model
hallucination yes U mod so um my experience with it my harble experience
with it is related to generative AI specifically right I think that model Hallucination is not really a concept
that would really use until until really that that spiked a lot with stable diffusion models and llms um so when it
comes to model hallucination I think what what we want to to be wary of or what we want to have in mind back to
what we said a bit earlier is explainability on the one hand and reproducibility on the other uh so being
able false discoveries is inherent to science right there are always false discoveries the fact that other
scientists other researchers around the world can reproduce the experiments reproduce um the methodology and find
new results is really what correct the the course and finding making sure that
we push as an industry and as an academic Community for the availability
again of the models themselves their code but also the availability of data sets uh to reproduce and create new data
sets in the future I think would help alleviate a lot this concern um so that
that’s really my take on it when it comes to D dimensionality reduction I’m not sure I followed completely your
sorry your question but um there there depending on the use case depending on what you try to optimize especially
between the local and the global structure that you want to preserve uh one visual visualizing the data there
are different methods that that you may want to use um you know in in single Cellar and AQ map is really something
that is used a lot in other fields you know other types of dimensionality reductions are used so uh my experience
with it as well is is really to to focus on what do you want to retain as a
priority from the latent space and really optimize for that and trying trying different methods if
needed so in in the area Dimension reduction one of the answers always used to be PCA or or vae you know a
variational auto encoder where you try to compress information and then reconstruct it and conceptually in both
cases you’re asking what the variance of your data is and having a good representation of that variance so you can build it back in the autoencoder
case um there have been interesting advances recently that might help in this space where from understanding
correctly your concern is it’s a very large sparse and diverse feature space because if you’re thinking about medical
records you’re thinking about individuals we we have a lot of different spotty different information about each individual and that’s a
problem with classic issue in machine learning where you need to have the same features the same inputs for everybody
and so there were early versions there’s something called Deep patient um seven
years ago maybe that it was a vae for patients where you just have a fair amount of noise and you do like it’s
called a d raing a encoder where you inject random attack on the information
and you ask the model to fill in the gaps in the modern era we call that masking or self-supervised learning and
so you can do self-supervised learning in an autoencoder context to help reduce your data and have a stable
representation for a lot very different patients that have a lot of very different information you might also
think about and this came up actually yesterday for those who are here this idea of contrastive learning which is an
example where you look at pairs of of information about the same patient um or you take pairs of information about
different patients that had the same condition if that’s what you care about and you can begin to an odd in sort of a
self-referential way build up a better condensed representation also known as embedding or lateen space to do that so
that would be one way to think about dimensionality reduction in the case of hallucination absolutely I mean Hallucination is generative um otherwise
it’s just a bad prediction and kind of the answer to both of those is we’re scientists we go test it and so if you
believe a model is reliable enough to go test the predictions it’s an empirical answer to how we think about hallucination and based on what we see
you can do Active Learning feedback loops in order to improve the model where it makes the most
errors yeah thank you thanks hi uh my question is also in
association with the former two questions which were asked about the difficult problems in biology right so
it looks like once llms came in everybody took out like the simpler ones and they’re also important because we
haven’t really solved like docking based predictions and stuff like that but what is the next froner of biology beyond
that and what is stopping us from going there is it is it something that at the level of like multiple models talking to
each other or is it simply the compute power and if you could talk to like building realistically teaching a
machine biochemistry where are we uh from from there
[Music] ah okay
so in the one hand there’s the the relatively simple question of hey when
are we going to be able to actually make models that can predict if this will be a good drug even if we already know what it is and we already have it in the
pocket and even then and the issue is training data it has just been terrible
and we’re so limited of having enough data to train models at this scale even even for really simple things say I want
to train a neural network to recapitulate something we know how to do because it’s an existing docking program
that doesn’t give us the right answer but at least there’s a formula in what we’ve seen is most state-of-the-art
models need 80,000 or more training examples to recapitulate a simple sum of log energy style docking function which
is for sure way simpler than reality is so we need way more than that orders of magnitude more than 80,000 training
examples we’re not yet at a place where it feels like we can get that from experimental binding affinity and so
maybe an intermediate solution would be some of the things called ABF absolute binding for energy perturbation
calculations problem is you can only do a few of those calculations per day right now using classical methods and so
there’s this problem then of we really want this as a way to speed things up we want to make sure we can do it in a
generalizable way we don’t have enough data the other I so so that is kind of
where we’re at and I think we just need a huge amount more labels in that space realistically um there’s ways we can
cheat and try to get ourselves to kind of build up a foundation using these these techniques we talked about that are Foundation model based that are
unsupervised that are self-supervised to maybe solve some of the basics of geometry first and then have the
learning process really focus on this really hard part of the question but there’s another sideways answer I want
to give to your whole question because I just answered a lot about binding affinities specifically which you mentioned at the end but wasn’t the
whole question I would actually say the biggest advances were about to see is
the final mile problem so many things um are were POS are possible right now so
many models are sitting on GitHub as we talked about and are shared and so few
people are able to use them or it’s the same people every time it’s the same Community it’s the computational people
who made the models who Honestly made models that weren’t solving the questions people care about more often
than not unless there’s that collaboration that relationship and so so that’s the final mile like in postal
delivery it’s really easy to get it to your ZIP code all the money and difficulty of a postal service is
getting it to your door so what is it necessary to get it to the bench and so my best answer is this people
collaboration question but if we can close some of that Gap and I know that’s a space that mlst is in and also the
application and use of it I think that’s where you’re going to see the biggest change if no single new model will developed today there’s so many uses of
the models we have right now that nobody is doing yet because of this friction
Enlighten very briefly um I think infrastructure actually is is critical
um we do so everything that you just mentioned is is brilliant I think what another Frontier that that we should
look forward to is is integrating the different modalities of biology right now uh we do have very
different lenses that we can use to look at biology at the DNA level at the RNA level at the spatial transcriptomics
level the prot omic level etc etc and putting all of them together creating really a model that um a mental model
and a biological mechanistic model that can put all of these layers together is really the next Frontier that requires
data curation and data collection to a level that is extremely challenging and that requires infrastructure that is
extremely you know difficult also to scale and to to put together so that’s really also something to look forward to
is the creation of these databases at scale and these intuitive tools for people to be able to access the data
without having to be computational people only right we need the biologists
we need the Noni and tech people to actually look at that
data well uh thank you everyone uh for those great
[Applause] discussion thank you now I invite you to
take a break and then we will gather again at 3 for the second panel about Ai
and Drug Discovery thank
you