← Back to news
Events

Health AI Symposium: San Francisco UCSF December 6-7, 2023

"HealthAI Symposium," a French American Innovation event at UCSF.

The Quantitative Biosciences Institute (QBI) at University of California at San Francisco and the Office for Science and Technology of the Embassy of France in the United States presented the “HealthAI Symposium,” a French American Innovation event at the University of California, San Francisco from December 6-7, 2023.

Supported by renowned research institutions, the health symposium brought together leaders, industry experts, and researchers from the United States and France to talk about AI applications for the health sciences.

MLtwist’s Audrey Smith joined Laurence Calzone (Research Engineer, Computational Systems Biology of Cancer, Institut Curie), Michael Keiser (Associate Professor, Institute for Neurodegenerative Diseases, University of California at San Francisco), and Hocine Lourdani (Group Product Manager, Deepcell) in the Artificial Intelligence for Cell Biology panel discussion, moderated by Moderated by Lisiena Hysenaj of Parvus Therapeutics.

Check out the full panel discussion (Day 1 of 2).

hello everyone welcome I’m jacen fabus the Chief Operating Officer for the quantitative biosciences Institute qbi

and I’d like to give you a warm welcome to today’s event uh today we’re kicking up our third joint event with the

scientific Department of the embassy of France and specifically the French

Consulate of San Francisco uh on the topic of research and health our first joint event was in

2020 it was virtual due to the pandemic and I’d like to say that we were slightly ahead of our time with the

topic being AI big data and health in 2021 it was still virtual um it was a

multi-day event titled reshaping science and health each year has been a great

success but I’m happy to welcome you to our first in-person event today in 2023

Health Ai and with this I’ll now pass the mic to our partner the French

Consulate of San Francisco represented by Emmanuel poak vour the atache for

Science and

technology good afternoon everyone um I’m Emanuel P I’m the ATT for Science

and Technology at the at the Consulate General of France in San Francisco it’s my great pleasure today uh to be your

co-host on France’s behalf I would like to thank our longtime partner qbi and

jacn for uh hosting us and co-organizing this event with us which is a French American event so you

will see a lot of French or French speaking people around and that’s normal I don’t think I need to convince

you that the topic of today’s Symposium is of utmost importance we’re proud that

France can contribute to these advances hand inand with a renowned partner like qbi I’ll try to keep the speech as short

as possible but I have a few things to say um one of our main missions at the French office for Science and Technology

of the embassy of France uh along with all the scientific attaches in Atlanta

Boston uh houon Chicago Los Angeles and Washington um our objective is to Foster

scientific collaboration between French and US researchers and support students

and research Mobility between our two countries especially at the doctoral level or early stage of the carea it’s

not just by chance that the collaboration between France and the US is is one of the strongest and the most

enduring our two countries do have a lot to share in terms of knowledge and knowhow in all aspects of society

including science and research we also value the supportive

entrepreneurial entrepreneurial spirit that exists in the tech world and which

is particularly strong in the Silicon valet as some of you may know so incubators startups spin-off companies

and large Industries are the ones that truly turn the most mindblowing Research

into a reality that will impact and improve the lives of the greatest number we could not be more thankful to

see French and American actors of innovation come together in an event like this one and none of these would be

achiev none of this would be achievable without their kind participation so I really want to thank them

truthfully my final message for today is that this Symposium is a good example of

what what we can achieve with all of your support uh as well with the French and

the Fran ofile community so it’s only one example among many initiatives

supported by the French Consulate on an everyday basis our goal our job is to listen to to talk about

the scientific and technological subjects which are of crucial importance to you and to both our countries and the

future of Healthcare in the AI era is definitely one of them we aim to create

dedicated spaces in communities where scientists can meet share major scientific results and insights that

will impact the research the economy and the society significantly there are many ways in

which you can support our actions whether you are French or not so please

come and talk to us during these two days learn more about the things that we do and the many ways that you can get

involved and just a quick word to maybe mention the Photo Expo that you you may

have seen when you came in this room which was one of the initiatives of the office for Science and Technology last

summer and we actually had some laurates uh who won some prizes in this photo competition so next year please

participate without further Ado I leave the floor to Tanya CM who will also

introduce our first panel discussion and I wish you all a very fruitful and enjoyable Symposium thank you

yeah good afternoon um and welcome from me um as well I am Tanya corta I am a

professor here at UCSF in bi engineering and therapeutic Sciences I’m also the vice dean of research for the School of

Pharmacy and personally extremely excited about the Confluence of AI and

Discovery Science reaching all the way from really fundamental curiosity to to

impacts on our everyday life and so to me some of the most exciting advances

happen at the intersection of Science and Technology of the intersection between different disciplines in science

and at the intersection um between technology science and healthc care and

so I think that’s why to me um what we’re witnessing um all over the world

but also what we’ve been discussing in the last day and a half uh um in a sister Symposium on AI across biological

scales and what we are continuing to discuss today in the panel discussions um will be really at this exciting

Confluence um of advances in computer science in cell

biology in drug Discovery in health care and in clinical applications and so I’m

very excited to hear from experts in this area um how they are viewing all of

these um developments and what they can inform us on what the future holds in

these areas and so with this I’m going to introduce um our first um panel

moderator um lisana heai to introduce her panel on

um cell biology and Ai and also the panelists welcome thank you very much

so the first panel will be AI biology and we’ll have four

speakers the first one it’s Michael Kaiser from UCSF he’s an assistant professor at UCSF I think this

microphone is not working okay yeah

okay and the second speaker is uh Laurence from um Kon from she’s coming

from France from Institute C and then we have oin Lani that is an engineer on AI

and O Smith from ml twist so every one of the speakers will present itself uh

for three minutes and then we’ll have a panel discussion together so maybe Michael you can

start

all right maybe we’ll put this all right well thank you so I’m an

associate professor here and my background is originally in computer science and over years I actually came

into this area through uh through UCSF itself I was once upon a time a bioinformatics graduate student here and

so the lab and we’re giving very brief ENT um our lab focuses um on three major

application domains in bio medicine and biological research for us Ai and ml

questions are at the center of it and so we look at molecules and we spoke a little bit about that yesterday about

perturbations and here when I think of a perturbation it could be a molecule it could be an environmental Factor it

could be a functional genomic change to a biological system and so some of the areas we’ve been interested in have been

cell biology as relates to this panel also whole organisms and collaborations

in zebra fish and looking at neuroactive activity and then Imaging Imaging can

span so many things in our area of Interest here I’m showing some examples of collaborations in the space of

digital pathology where we work on becoming a force multiplier or a way to help Pathologists operate at a scale

resolution and generalizability that otherwise would be impossible to do as a

human and so we were asked to show one or two examples I’m going to take the easy route and show two that were from

the presentation I gave yesterday the first one um is a generative AI model

where we were ingesting images these are microscopy images from cellular screens

for a neurodegenerative disease phenotype specifically the tow protein

and at the left what you have are two fi are two different channels of the same field of view of these cells

at the top you see yfp fused tow protein and at the bottom we have some nuclei

now as I mentioned at the time one of the big problems with live cell Imaging to do drug Discovery and other screens

at scale of hundreds of thousands or more perturbations is we can’t put antibodies in that system we can’t run

it in a fix cell format it doesn’t scale and so if you’re doing this in live cell and we also look at different time

points we’re stuck with only seeing all of the tow what we really want to see though is the misfolded tow or the

Tangled towel this is the tow that builds up pathogenetically and we see it in the brains of patients who have

passed away from Alzheimer’s disease and that’s only some of it and so what we were able to do was train a deep

learning method that could take data from screens that had been done years before and ingesting those two input

images on the left generate what we wished we would have had all those years ago and specifically focus on the tanks

Town happy to talk about that and then the other sort of result that

illustrates area of interest for us is a second is a second generative approach and so this is a generative diffusion

model and what we’re able to do is generate fragment by fragment piece by

piece molecules within protein pockets and so combining these two approaches

you can imagine a whole pipeline of testing experimentation and assessment

in order to move forward understanding cellular diseases and Drug Discovery

thank you thank you very much and so now we are still in Academia and we’ll have

someone that comes from France L will introduce what she’s doing and who is

she thank you so my slide are not as beautiful as as what I hoped but um so

my name is Laurence Calzone I’m a researcher in the Currie Institute um I

coordinate together with Emanuel bario with in the audience the group of computational systems biology of

cancer um I’m a mathematical modeler and with the focus of uh on cancer

applications and what do I do what am I interested in is to build mechanistic model that describe the biology of the

disease and in our case in particular cancer but not only with the purpose to

understand the impact of of the alteration suggest optimal drugs and predict drugs response or drug synergies

in a patient specific manner um so where do I come from so I’m

a mathematician by training uh I uh I I started doing uh chemical kinetics in

particular and then I move to biology through Theory and I did a PhD in theoretical biology here in Virginia

Tech not here but in in the US uh so my first interest were to do some dynamical

modeling uh of uh of the cellular processes and uh and then when I moved to K uh the

the type of data we were dealing with could not be applied to ordinary differential equations really so this is

when we developed some some other formalism uh based on on both um

approaches but what we we really develop in the group are models of tumor Evolution and expansion uh taking into

account space and time so with have a a big big data and and a lot of parameters that we don’t have uh we also developed

some tools to model individual cell populations and we want to integrate of course all the omx data that we can into

our models so um I’m really interested in

mechanistic models and how to make the link with artificial intelligence models you know the we try to uh explain really

the type of results we get with AI models and we can do that from two perspectives first as an input to the

models so we learn from the data to build our mechanistic models of course with the stratification of patients but

also help us infer the models that we build both the networks and the the mathematical equations from this data we

want to integrate multimodel uh data as model parameters so again learn the parameters from this U from this data

and of extract features from the data which PA Pathways or which genes really separate the response to a to a

treatment or not and then when we have this model we are also interested in in in uh um AI to um improve the models and

how do we do it uh we search for optimal parameter sets so we have a lot of parameters that we need to infer and

this is this is really difficult to do it by hand and of course we expect to to use um uh AI models like surrogate

models to mimic uh the complex and Compu

computational heavy models that we build and we also use AI to generate synthetic

data for model optimization we don’t have sometimes we don’t have enough data and this is one way to augment the data

so we want to make the link between AI results and model outputs and really fill the gap between mechanistic models

and Ai and I think I will try to give this perspective in uh in the response to uh to the the questions that we will

we will talk about um the point is to explain the content of content of the black box as much as

possible and and I was also asked to give two uh two um um projects on which

we work in the lab the first one is uh is called  is um the point is you know it’s the holy Grail to try to find

some biomarkers for a response prediction to to some treatments and here it’s about immunotherapy and we use

multiple data types to uh increase the the power to be predictive and explain

as much as possible uh why these patients respond or not and this of course open the road to special modeling

which brings me to the second project uh which is my expertise is spatial

modeling of tumor Invasion we model both intra and extracellular interactions um to uh to simulate in

silico treatments so we start from data we infer the model we infer the parameters and you can see here an

example of of a tumor and and invasive um that starts invading the the tissue

and then we can play with these models and search for optimal drugs in silico that we can then test um in uh in

experiments and I will finish here okay thank you now

it’s oin it’s a farm D that is going to explain us more what he has worked for

some companies in the Silicon Valley on AI so he will introduce

himself sure hi everyone thank you every thank you for inviting me and thank you

for being in the room today really appreciate it quick int introduction by myself my name is Jos Lord I am a farm D

indeed By by Train um started I am originally from France which I wouldn’t mention uh usually but feels relevant

today i started my career in in France actually in Pharma in biotech then went

to business school moved to the Bay Area a bit more than 10 years ago and since then mostly spent my time building AI

products and so I I’ll try to bring that perspective of a practitioner today and um and somebody who has led the

development and the commercialization of AI products in different Industries in

part Healthcare and uh and Life Sciences Industries so a couple of examples uh I

I extracted here uh is the work of uh I’ve conducted with the teams at arteries

and Rapid AI where the goal was really to leverage AI in a clinical setting um

to build software medical devices to support uh clinicians in their diagnosis

and treatment decisions so very specific set of challenges are related to that

it’s really about the measurement of certain clinical uh factors the

detection of certain pathologies and that’s a very specific type of of challenges that I’m sure will come up in

our discussion most recently I I’ve worked at Deep Cell where we took a very

different approach and Deep Cell is a company focused on Research especially at the Single Cell single cell research

and we applied AI to study the morphology of cells so we built an in end platform um able to take images of

cells in flow and extract High dimensional information from these images and help researchers characterize

and make sorting decisions U using um using the outputs of our foundation

model so very different set of of of challenges here and so I extracted just a few thoughts that may color my

intervention in this panel very high level and I won’t take too much time one is um I’m a product person

so I’m really focusing on building uh products and solving concrete problems

as we know here these days there’s a lot of hype a lot of excitement and for good reason around AI um but beyond buzzword

beyond the hype um I tend to focus on identifying very specific concrete problems and focus on the impact that AI

can have on solving that problem uh with an emphasis on the inter interpretability and the usability of

these models because the best model in the world can be very is not very useful

if um the the end user cannot really use it in good conditions um the second

element I wanted I kind of touched on is that challenges in AI develop product development and commercialization

implementation are very different depending on the on the industry uh in particular as I mentioned clinical uh

environments versus researchers only research focused environments are present very different types of

challenges and that may uh color my intervention today and finally one point that I wanted to to insist on and

actually that resonates a little bit with what Lance I believe was was presenting is that one thing I’ve observed in um recently in this and in

my my time working with biologists and researchers is that there is a persistent cultural difference Ando

difference in approach between researchers in biology and Ai and the AI community at large and it focuses a lot

thought and my experience in explainability and in whether or not a model actually describes a mechanism and

of course this is a complex topic so I W this is way too high level but I want to emphasize here that sometimes that qu

that question is very important and of course a model should be explainable and at the same time the value of a model

can be in its predictive power alone that happens and an example that I didn’t work on personally but I think

many of you in the room can relate to or or know of is Alpha fold for instance this model trained by Google Deep Mind

to predict the 3D folding structure of a protein based on its sequence of amino acids there is work on going on

explaining exactly how the model works and how it’s making its predictions but the value of the model in unlocking

avenues for research um just by being that potent and that that efficient in

predicting the structure is is already huge and sometimes explainability is not um the the only

aspect that we should focus on so I’ll stop here I took already too much time thank you everyone and I look forward to the to the

discussion thank you and so now we continue with other four panel and she’s

alre she’s a woman founder and she has build this company that is on AI

Services she will explain more about it hi everyone thank you for having me

so um two facts about me I’m French as you can probably be here uh and the

second one I’m the only non-scientist person in this panel so very uh very happy to be here very excited to be here

but I might offer another uh another point of view um so um I founded ml twist uh with my

co-founder three years ago and the idea is really uh to help soft a very um very

boring problem actually uh but that is like very persistent when you uh start training a model is um on uh the fact

that there are still like manual task that still need to happen when you need to create data processing or data

augmentation and when you try to not only build data pipelines for certain model but also maintain it because it

can break over and over again and that type of work actually um is taking 80%

of the time of any data scientist which is which is a lot and most of the time if you talk with data scientists what

they want to do is spend that time training a model not doing that genitori work that we’re happy to take

over um if you look about like the yellow shun you’re going to see that as

of right now a lot of um Technologies happen a lot of technologies have like raised over the past decade uh but what

really is not really like you cannot see happening is that you need to build the

data pipeline between all those yellow box to make um to train your model and that’s something that still not have

been like really solved as of today and that’s what we focus on if you look at the work uh the manual

work if you have a technical person on hand in your company or in your uh in your work um in in Academia you’re going

to have to go through all those different steps uh to get your data ready for a machine learning model to um

ingest the data and get trained on the data that can take up uh like several

weeks as you can see um and we build a platform that can auto generate data

pipeline um and all the uh purple uh steps that you can see are done

automatically by our platform which um um help people like you or data

scientist to just focus on you know like the exciting stuff which is uh working

on the accuracy of the data or just training your model and and checking the performance of the model

um as you can see we have been uh lucky enough to partner and uh have also great

customers that we allowed to talk about so we are working with s National Laboratory but also UC Davis health um

we are also working in uh with Stanford uh HAI  department and um partnering with

Cap Gemini and and other uh names on on this slide um and so we have been able

to work work on a lot of different data types uh from what I was talking about

like images but also videos um it can be also text and it can be also dicom uh we

have actually worked on daos which is a spin-off of dicom um and that’s about it

that’s us in a nutshell oh you can you can see like a bit of our accomplishments too um

but yeah that’s that’s about it thank you thank you

so okay thank you very much everyone so as we saw in this panel we have uh

people from Academia that work really on building this model and making sure that we’ll have some discoveries from these

models and then people from industry that are making sure to create a product that can be used for everyone so now the

question I have it’s what is one of the

most impressive discoveries that we have had in the field of AI and cell biology the

recent years that we could not have had just with humans alone anyone wants to answer this

question maybe can you hear me yes maybe I’ll get started just because I just

reuse the example I took actually I mean in my presentation um I think Alpha fold

is is one of the most impressive developments that we’ve seen lately um

uh the it’s interesting to see that it’s a collaboration with Academia initially with emble um in Europe to build the

database and then the development of the model um you know was conducted by

Google deep mine U to your question Lana so the the problem of protein folding

and the 3D structure the prediction of a 3D structure of a protein is a well-known problem it’s been worked on for decades it’s been worked on

tirelessly by a number of groups and uh I’m sure many of you are in the room are aware of it so I won’t spend too much

time of it but the on on it but the level of performance that was reached recently in 2018 with the first version

and then 2020 um is just something that we couldn’t have um reached without the

latest development of AI um now you know I when we say would would not be

achievable by humans I’m always a little bit you know not sure what we mean we always use mathematical modeling we

always use different tools and these are created by humans and AI are also created by humans so I would still you

know put that on on the credit of humanity uh but this is I think a very

solid example thank you so another thing this is like maybe in the mind of everyone it’s we have seen all the

Sci-Fi movies where we see that okay something it’s taken by a malicious

attack and some criminals took over our mind that in this case it will be our

machine learning and I was thinking this when I heard the talk of Michael where you were talking about deep learning uh

techniques that you’re using in the lab and how you can determine it’s like a towel Protein that’s misfolded or all

this so how you make sure that someone will haug your like how how are we sure

that what do you do to protect your system and someone imagine someone comes

Haws your system and then your students or PhD students or post dog will use a

model that is not right and will claim discoveries that are

false ah well so this is it’s an interesting question too because in AI um there’s so

much of a history of releasing our code this is crucial if you want to publish

at a computer science conference um if you want to do any of these things you have to make what you’re doing available

and we’ve seen questions about this with things like chat PT and the question of guard rails or improper use but when we

think about drug Discovery um there have been papers where people have asked what about old techniques what about things

we’ve known for decades computationally how hard is it to misuse a model and that has always been

possible it hasn’t been the computation that’s kept the the computational barrier usually that’s kept someone for

making a very bad drug it’s been the synthesis the construction all of the other steps where the rubber meets the

road where there’s actually a lot of oversight and a lot of institutional support and that is one way that the

models themselves are less at the center of that question um but what you are

afraid on it’s not the the model it’s like this malicious attack don’t exist it just like we are scared of it I am

less afraid of models than I would be about how people choose to use things okay and so for me I’d always take that

answer to our organization do the universities the companies oversight

accountability that would be the place that I would ask because I can because the flip side of this is I remember

years and years and years ago I was um a student and someone was presenting on

some new technology or result a professor who was visiting and it wasn’t

published yet they were hoping to submit it to a really high-end journal and someone said but aren’t you afraid that

you’re by telling us all about how to do do this you know someone’s going to scoop you or use it in some way you don’t like and he said young man um I

would be thrilled if someone would try to steal my work okay great so thank you

for your answer yeah no just one thing about I mean in Academia we share as much as possible so we we put on GitHub

everything we can whenever we can to make sure that also we get the paternity of of what we develop but also uh like

you said we we hope that people can use it and then make it better and then we all work for the same purpose okay thank

you so one question I’m having now because I was searching a little bit

about ml twist and your collaborations and so that I thought ml twist is a

service and it’s used only by companies but I saw in your website that you collaborate with Stanford and Berkeley

and I was pretty surprised about it can you explain how you help this big

academic institutions that generate a lot of discoveries in AI to improve

their data sure so um we have like several ways to work

with them or to partner with them uh one of them is actually using uh the

students at those universities that like our medical experts for instance and who

can help us uh do some labeling uh for some some models uh we have other things

like you said like for Berkeley labs it was a different because we were working on um uh a grant uh for the Department

of energy and we were able to collaborate with them on a plan to win

actually the phase one of those uh of the that beer um and every every time so

we also working with 10 for hii and that’s something else like they are working on uh creating some models some

llms uh internally and we are generating the data for them so um that’s that’s

like we are very very proud of that uh but like every every time we talk with an Academia it’s like going to be a

different need uh but we’re happy to explore any way we can partner with uh Academia okay thank you so basically we

we saw that Michael said that everything the discoveries are done and the models are not misus as people use them

properly and then we have some people to help label well the products in order to

generate even more discoveries now as a biologist or immunologist that I am the

question I am asking it’s I’ve heard all this different models and companies that

exist is there any way I can learn because I don’t know how to code I can learn how to use or where I can be

informed about this new technologies that are coming

out um it really

so I guess I’m coming at this question from a again product perspective I tend

to believe that not everybody needs to be trained on the technical aspects of AI and that is the role of companies

largely but also of course the contribution of the academic Community to come up with tools that are very easy

and intuitive to interact with that in the background do very complex things and use a lot of code and use a lot of

very complex mathematical concept cep but from a user perspective just you know kind of Shield all of that from

users um I don’t see a very I may being correct and you may have a different

opinion I don’t see a very strong argument for AI to be any different than software engineering in general for

instance we’re not asking everybody to be able to code as you just said uh in research and I don’t see why that would

be different from AI um I think the argument is a little bit different when you think about students however uh when

we think about students in the fields of of biology I think the the the intersection of computational biology Ai

and my biology uh in general is reaching such a level that there is probably a

better argument there to say that the Next Generation in general should touch at least to some degree uh AI training

uh but it’s a field that’s evolving so fast anyway that you know it’s something that is a very complex convers you know topic and I don’t have the the expertise

on it thank you no I would say that uh the key for any project is collaboration so as a

cell biologist you don’t need to know AI but you need to be surrounded by people that can help you and I talk the same

language as you yeah so this is the question it’s like where do I find because every time I go Ai and then you

have this courses I’m like oh I’m lost what I do so the question was like where can I find this Char GDP for example was

all over and everyone was talking about it so I started to use it and it was easy and save my life but where is there

any website or something you well I I would emphasize that the collaborate point because that’s what I was thinking too because actually I don’t think chbt

is the right metaphor for how we often want to use AI for high stakes problems

because instead I I think of you know someone who uses the centrifuge without

training or the fact machine early on and now it’s gummed up either you’re

working with someone to at least get yourself started or there’s a core facility or there’s a way to get the

training okay and the problem with AI is let frequently that we would gum up a

machine or do something to the rotor it’s more commonly that it would appear to work in a way that was misleading or

not useful because it is a pattern recognition thing it will always find patterns and exploit them so collaborate

collaborate thank you and so uh I’m wondering just like you are all open to

have a contact from someone from industry and build a new model or help them on a this is this

is great I think it’s a message we have to the show so I think we are I don’t

know about the questions or I have five moment okay perfect so you’ll tell me

whenever yeah I have more questions so another question that I had it’s okay

we’re in the same space we need to be trained do you think like a training of AI like how to use Ai and not to misuse

it should be mandatory in every company now or in every lab like the safety

training um I’m not sure exactly how it would look like but I think that it

would be very important to start educating people on the good and the bad

um starting by data bias uh which is like something that is like as human we

we all have biases but if we are of it if you let people know that they need to

pay attention to the data they’re going to be selecting for their model um then there might be a chance that we’re going

to reduce the amount of bias happening in the model um and then to go back to your question where you were talking

about like learning how to code on on what it means to do Ai and so on I think

it’s actually important to keep that separate just because you’re going to have people from different backgrounds

collaborating and they will fight also biases that way if you have everyone you

know educating the same way learning the same way in the end you’re going to have a lot of biases okay thank

you just again just to say that I mean in the company you need a wide range of

expertise again so you need to have people dedicated to that and the other ones knowing the limitations of of the

results and how to use these results is really key I think this if we have to all of us if we have to learn something

is the liit ations and the biases to be clear about the biases

yeah and just one thing to say about it um because of the strength of all the

different expertises and opinions we don’t all need to be programmers or developers of AI what will happen more

and more is these tools in different ways make their ways into the workplace into the research environment all of

these and I worry that it seems like when you read articles when you think

about and you talk to people who have questions outside of this space it seems as though it’s categorical do I trust AI

or do I not trust AI um do I use a prediction or not but when you think

about things like chat GPT which I I’d like to you bringing up earlier we are used to speaking and writing and we can

critique it we can look at it we know sometimes when it’s wrong and if you try to look for that you find it more and

more and so I would say the training that would make the most sense is how to be a critical user of it yeah and what

techniques to have to test whether this thing is telling me what it so confidently is

presenting I completely agree with that point and to take it one step further I

think one one thing that we need to put on the agenda even more is the question

of Open Source uh in the industry in particular because indeed in Academia and we mentioned it earlier there is

this culture and this requirement to share your your results to share your

code to share your data and in the industry it’s not always the case and for good reasons of course if you

invested a lot of money in Training Systems you want to protect your IP you want to protect your first mover

advantage in the market and completely understand that but we need to find ways to still have enough data available and

enough source code available for researchers in Academia or or not to

actually validate these models and and um and test them and be critical about

their performance thank you so basically I’m more optimistic now on the use of

the AI and the models so now just to talk about cell biology in general what

is as as an immunologist I’m more like in characterizing the cell types and

it’s basically what you have done a little bit in Deep Cell so like will we be a moment we will just recognize cells

by uh only one feature just for example

Imaging or we’ll show could still do our flowetry with 300 different colors and

then analyze by ourselves and how how far is this I’ve heard also like in in

the same space I’ve heard also about wet lab so if you can explain this and where

we are in cell biology right now and what the Yeah I can draw from my recent

experience at Deep Cell to at least answer the first part of your question um so we are at the point point right

now with the the performance of AI models on Cell Imaging and we can really extract High dimensional information in

a way that is consistent an improper way to to say that would be to sequence the morphology of cells and at at scale

right and that’s completely new and that really challenges or or subverts let’s

say the the existing methods in facts and fetometry um where I think uh it’s very

interesting is that currently the way to characterize cell is based on the taxonomy in biology that is in evolution

but still very much rooted in surface markers cell markers in general and

getting out of that tonomy initially is a little bit different when you start looking at phenotype and that’s the the

next Frontier is really looking at phenotype for itself and really leverage that as a new omic basically uh type of

of data um and the next frti after that I think is multimodel you know research where you integrate that data with

multiomic information from different sources um so that’s that’s I think how you you get to a holistic view of a cell

that takes very different um you know lenses to analyze it did I did I answer your question I I think so maybe Michael

has something to add about I well I very much agree with what you’re saying and what I was also kind of reflecting on

and thinking about this is you ask kind of when we’re there um and I remember um

1965 a paper on strong inference quote from it is the measure of a method is its use and so we’re talking about

methods and technology and platforms AI or otherwise or combined in this hybrid system for what use and so for some uses

we might already be there right now and for some things we we cannot see them by

Imaging without the right marker and we don’t know that yet and so there’s this interest of multimodality how many

things we can bring together without breaking the bank on the experiment in the first place yeah and so I would

actually turn the question back to you and ask for what use um and what is sufficient and I I was thinking more

than Discovery and cell biology for example could be used in the biomicro space so for example I would imagine an

AI um like model that could help me

realize what cell types I have based on some phenotypes that I see on the

microscope or what can be the molecular pathway that

are affected on this phenotypes and I saw that you are doing something like

this maybe in the lab to link this maybe also in Pharmacology and this is what

how you see this feel and what are the discoveries in

there so this could be an hour conversation I’m sure others you know have thoughts

on to in there’s big differences between supervised and

unsupervised we’ll go and see what’s there based on the of information we are

already toed up to collect say images ands compared to I have a really

specific need a priority bringing in my human expertise I know that none of those image channels are are likely to

get the job done and if they did I’m actually kind of that would actually be a good counteract test to make sure I

couldn’t get that out that maybe I

need so would best way can

is that’s what we as scientists

allim as as

possible any have question I think when

I R te so how how this

works

resources question canor from cyber and so in cyber you never know

your computer all the best

practices down ACC don’t want never and

that’s what about else C

guelin scient

unless unless you a lot of to unless you scr dat and the model

itap

tocce controls now we needer controls because we will try to do what we ask

and that’s the difference and so that’s what in Cy

sec

um so not that M because we are generating the data we’re not generating the models but that’s definitely

something that needs to happen uh for our customers so we work with defense

for instance and that’s definitely um mandatory to to have like

some ways of trying to break the model um but you can do it yourself you can just go on chat GPT and and then ask

questions again and again and you’re going to see at one point it’s going to break and you’re going to even see some

biases appearing like very quickly on what the pilot is it is a man or is it a

woman um or nurse um and and you’re going to see that this is actually all

coming from how the model was trained and what data was was you know used to

generate the CH GPT so that’s very important and that goes back to what you

were saying Michael that we need to be able to understand that whatever we see

that is AI stamped we need to criticize it we need to be able to look at it with

an eye of hey is it real or is it bias or can I really trust what what it does

or what it says um and that’s like something that should be done even uh at the school level like with our young

kids uh because with like everything that is gener it right now we don’t know

what is true and we don’t know what is uh not true and that’s very important to train kids and students to keep that eye

uh out for for that

do no I agree but this is a scientific reasoning right you you want to see how far you can go and and find the

weaknesses of of your models and this is something we do uh even when we read an article we don’t we don’t take it as as

it is we want to question it and I guess this is what we need to keep in mind with AI models they’re not the truth

they they get closer to what we want and they answer a very specific question not Universal questions thank

so now as as we heard we not be scared of

thei they work in the same way as we work for experiments so they check every

step and and also they’re all working in a collaborative way so we should be afraid

to go and ask them if they

industry and

then and everyone if you have any questions for theel could should line

up there will be a microphone here so you can line up and ask

questions I start thank you so much for the panel uh so I have a question

because I see a lot of like academic or pure Tech kind of players how do you

work with the Biotech Industry how are your Solutions implemented do you feel

that there’s any bottleneck um from the pharmacetical industry to like adopt all these great

Technologies right well so one of the best answers I can give is the second result I showed with the small molecules

is actually a joint development with Genentech and so jentech um Partners came to us and we worked on a close

collaboration for more than a year here where there’s both Mutual support the IP

is worked out ahead of time by the university and Genentech and by setting that road up nicely then what we’re able

to do is get real feedback because one of the things that stood out to me in that collaboration was we made the

molecules they looked pretty good but our our in the early version and our collaborators said yeah but that looks

really strained like super strained in practice we’ve tried a whole bunch of public techniques and and the molecule

it has an that’s just completely UN unworkable in practice and so it’s

having at least a champion on each side and people working together directly

would be my answer to you on that no I can just say that we are in

France we are promoting a lot this uh this PhD program where you have industry and uh and Academia working together so

we are really uh try to uh teach students to embrace space industry and

uh and think about all the the positive things you can you can have like you know access to clusters access to money

to experiments that we may not have in

Academia one more thing I would say is that a lot of the startups that are trying to bring innovation in the field

of AI are often spin-offs actually of academic institutions and so that in Industry uh Academia collaboration is

almost you know inherent to a lot of the Innovation that’s ongoing now your question was more towards Pharma and

biotech at a large scale I think the Pharma and biote can be a difficult

industry to penetrate initially when you’re startup they’re looking for usually a lot of validation a lot of

very early data however I think in the field of drug Discovery in particular AI applied to drug Discovery you see a lot

of different Partnerships that are actually ongoing um Asen with benevolent AI for instance is an example okin with

sop sop is and another one there a number of of of Partnerships of the of that nature so I think we’re going in

that direction more and more and fora companies realize that they can embed uh AI um models uh and products very early

in their development so you you think that in five years from now there will be more and more drugs that are in phase

two phase three that have been like discovered thanks to AI definitely yeah

this is already happening they’re they’re already in the pipeline but very early so we’ll have to time will tell

thank you thank you I might because uh just a

quick comment before there is a lot of ethical consideration and I will be hosting the ethics and AI next week and

I see you know like it’s very tomorrow and I it’s very relevant but anyway my question was more as a sale biologist

and we we we generate more and more like big data set like omx data and and all of that and we build model and test

those models um but kind of on the other flip side where you were mentioning like

usually we want like an end product and the end prediction with AI but can we

flip things around the bit and use AI to help identify Gap in our knowledge and

how should we build framework to help actually scientists trying to make sense of those very large data set that are

very complicated and have a lot of connection and help hypothesis

generation if you have any opinion on that

all right I think there’s two levels to answer first answer the kind of glib one but one that we want to get to working

is all of the explainable AI techniques interpretable methods saliency mapping

oclusion or even building models that operate through rationale layers where you’re having it solve multiple tasks as

a series of steps that are fully differentiable all the way through and if the model works or does not work

through that that information gate then that tells you something because you can change what information the task has to

pass through and so that’s one category of answer a a technological answer an

architecture style answer in practice much of the saleny mapping methods

integrated gradients guided grab cam either bring in their own artifacts or smear out signal and can only tell you

really basic things like is it looking at the dog or the background but trying

to understand about the ear the know is iffy at that resolution and in molecules goodness gracious me we don’t even have

as much intuition and so that answer is where the field is trying to go methods

wise the flip side answer is is a is a simple one and it’s a logical answer instead which is this is just one

technique and if anything use very simple models wherever you can standed

up swap in and out statistical models and other models and ask whether the bang for the buck is worth it to bring

in something that we have less of an understanding of and so that goes back to being scientists and the fact that

all of us here in the room have some have agency in making that type of decision without having to rely on on

hoping that a model that learns for reasons we don’t know can tell us something new that we didn’t know to look for so it’s considering it as parts

and swapping the parts no I I I would just say that maybe

uh one way to look at it is to use these models to validate an intuition that you may have so the question will be clear

it will not identify gaps but it will confirm or infirm uh an an intuition

that you may have one last thing we mentioned earlier

in the discussion the difference between supervised approaches and unsupervised approaches I think that plays a role in

your question when it comes to identifying gaps building taking an unsupervised approach to look for high

dimensional uh structure uh global and local structure in high dimensional space

using what is in many cases now known as Foundation models so models that are looking very

generally at at a field and then can be trained more specifically for a given application is an is an approach that’s

very relevant you can look you can really that that’s what I would say just from the the the ml training approach

just being unsupervised or self-supervised yeah helps a lot thank you I like that you said

intuition that likely scientist are not going to be replaced

by hi so my little thesis was in proteomic and omic data at Stockholm

University I wonder One Thing how can you deal with the biophysical Pro

properties of a shape shifter proteins like Spike because we had a lot of

problem to read the cre data talking of a particular class of proteins that they

are able to change confirmation because of

pH the short answer is no one’s dealing with that yet the ways you could go about trying

to start would be to think about your representation and because when we start

a model that in order to do any deep learning or machine learning training we

have to make a decision first and this gets tol twist for instance and thinking about what our representation of our

information is not just the foul format but what is this this cloud of atoms

what is this as far as the model is going to ingest it by the time it sees it for the first time in the world and

so geometric neural networks might be good you can make graphs spatial graphs

in space and then who’s to say that a graph only has to operate in space it

can also go across time you can put edges together to have on embl so that’s something similar to like convolutional

approach but in multiple dimensions and now you’re convolving in non-cartesian space so that would be one way to do it

you can use attention as well um so there’s a lot of methodological answers oh but yeah I want just to tell you that

for desperation we started to look at possible binding sides for transcription

factors oh yeah well let’s continue this after because there’s a bunch of neat

methods that can be done with this thank you

good afternoon uh wonderful discussion for the planel my name is bushan I am a bi engineer and I develop bi materials

and cell therapy uh I’m new to um AI in B engineering so pardon my knife

question so I have two part question one is uh what’s the panel’s take on model

hallucination and how it impacts uh uh the the uh clinical data or clinical

output uh or in in output in cell biology as well and the second part of question second part of my question is

uh we are talking more about going high higher dimensional in terms of analyzing object or making uh new objects by the

by the means of generative AI uh but in terms of clinical translation since I work on translational side of a science

uh we want to have a simplistic system to be uh clinically translatable so what is the panel take on dimensionality

reduction uh going forward after we have a lot of higher dimensional data

in based on the nodding um the first part of your question sorry uh model

hallucination yes U mod so um my experience with it my harble experience

with it is related to generative AI specifically right I think that model Hallucination is not really a concept

that would really use until until really that that spiked a lot with stable diffusion models and llms um so when it

comes to model hallucination I think what what we want to to be wary of or what we want to have in mind back to

what we said a bit earlier is explainability on the one hand and reproducibility on the other uh so being

able false discoveries is inherent to science right there are always false discoveries the fact that other

scientists other researchers around the world can reproduce the experiments reproduce um the methodology and find

new results is really what correct the the course and finding making sure that

we push as an industry and as an academic Community for the availability

again of the models themselves their code but also the availability of data sets uh to reproduce and create new data

sets in the future I think would help alleviate a lot this concern um so that

that’s really my take on it when it comes to D dimensionality reduction I’m not sure I followed completely your

sorry your question but um there there depending on the use case depending on what you try to optimize especially

between the local and the global structure that you want to preserve uh one visual visualizing the data there

are different methods that that you may want to use um you know in in single Cellar and AQ map is really something

that is used a lot in other fields you know other types of dimensionality reductions are used so uh my experience

with it as well is is really to to focus on what do you want to retain as a

priority from the latent space and really optimize for that and trying trying different methods if

needed so in in the area Dimension reduction one of the answers always used to be PCA or or vae you know a

variational auto encoder where you try to compress information and then reconstruct it and conceptually in both

cases you’re asking what the variance of your data is and having a good representation of that variance so you can build it back in the autoencoder

case um there have been interesting advances recently that might help in this space where from understanding

correctly your concern is it’s a very large sparse and diverse feature space because if you’re thinking about medical

records you’re thinking about individuals we we have a lot of different spotty different information about each individual and that’s a

problem with classic issue in machine learning where you need to have the same features the same inputs for everybody

and so there were early versions there’s something called Deep patient um seven

years ago maybe that it was a vae for patients where you just have a fair amount of noise and you do like it’s

called a d raing a encoder where you inject random attack on the information

and you ask the model to fill in the gaps in the modern era we call that masking or self-supervised learning and

so you can do self-supervised learning in an autoencoder context to help reduce your data and have a stable

representation for a lot very different patients that have a lot of very different information you might also

think about and this came up actually yesterday for those who are here this idea of contrastive learning which is an

example where you look at pairs of of information about the same patient um or you take pairs of information about

different patients that had the same condition if that’s what you care about and you can begin to an odd in sort of a

self-referential way build up a better condensed representation also known as embedding or lateen space to do that so

that would be one way to think about dimensionality reduction in the case of hallucination absolutely I mean Hallucination is generative um otherwise

it’s just a bad prediction and kind of the answer to both of those is we’re scientists we go test it and so if you

believe a model is reliable enough to go test the predictions it’s an empirical answer to how we think about hallucination and based on what we see

you can do Active Learning feedback loops in order to improve the model where it makes the most

errors yeah thank you thanks hi uh my question is also in

association with the former two questions which were asked about the difficult problems in biology right so

it looks like once llms came in everybody took out like the simpler ones and they’re also important because we

haven’t really solved like docking based predictions and stuff like that but what is the next froner of biology beyond

that and what is stopping us from going there is it is it something that at the level of like multiple models talking to

each other or is it simply the compute power and if you could talk to like building realistically teaching a

machine biochemistry where are we uh from from there

[Music] ah okay

so in the one hand there’s the the relatively simple question of hey when

are we going to be able to actually make models that can predict if this will be a good drug even if we already know what it is and we already have it in the

pocket and even then and the issue is training data it has just been terrible

and we’re so limited of having enough data to train models at this scale even even for really simple things say I want

to train a neural network to recapitulate something we know how to do because it’s an existing docking program

that doesn’t give us the right answer but at least there’s a formula in what we’ve seen is most state-of-the-art

models need 80,000 or more training examples to recapitulate a simple sum of log energy style docking function which

is for sure way simpler than reality is so we need way more than that orders of magnitude more than 80,000 training

examples we’re not yet at a place where it feels like we can get that from experimental binding affinity and so

maybe an intermediate solution would be some of the things called ABF absolute binding for energy perturbation

calculations problem is you can only do a few of those calculations per day right now using classical methods and so

there’s this problem then of we really want this as a way to speed things up we want to make sure we can do it in a

generalizable way we don’t have enough data the other I so so that is kind of

where we’re at and I think we just need a huge amount more labels in that space realistically um there’s ways we can

cheat and try to get ourselves to kind of build up a foundation using these these techniques we talked about that are Foundation model based that are

unsupervised that are self-supervised to maybe solve some of the basics of geometry first and then have the

learning process really focus on this really hard part of the question but there’s another sideways answer I want

to give to your whole question because I just answered a lot about binding affinities specifically which you mentioned at the end but wasn’t the

whole question I would actually say the biggest advances were about to see is

the final mile problem so many things um are were POS are possible right now so

many models are sitting on GitHub as we talked about and are shared and so few

people are able to use them or it’s the same people every time it’s the same Community it’s the computational people

who made the models who Honestly made models that weren’t solving the questions people care about more often

than not unless there’s that collaboration that relationship and so so that’s the final mile like in postal

delivery it’s really easy to get it to your ZIP code all the money and difficulty of a postal service is

getting it to your door so what is it necessary to get it to the bench and so my best answer is this people

collaboration question but if we can close some of that Gap and I know that’s a space that mlst is in and also the

application and use of it I think that’s where you’re going to see the biggest change if no single new model will developed today there’s so many uses of

the models we have right now that nobody is doing yet because of this friction

Enlighten very briefly um I think infrastructure actually is is critical

um we do so everything that you just mentioned is is brilliant I think what another Frontier that that we should

look forward to is is integrating the different modalities of biology right now uh we do have very

different lenses that we can use to look at biology at the DNA level at the RNA level at the spatial transcriptomics

level the prot omic level etc etc and putting all of them together creating really a model that um a mental model

and a biological mechanistic model that can put all of these layers together is really the next Frontier that requires

data curation and data collection to a level that is extremely challenging and that requires infrastructure that is

extremely you know difficult also to scale and to to put together so that’s really also something to look forward to

is the creation of these databases at scale and these intuitive tools for people to be able to access the data

without having to be computational people only right we need the biologists

we need the Noni and tech people to actually look at that

data well uh thank you everyone uh for those great

[Applause] discussion thank you now I invite you to

take a break and then we will gather again at 3 for the second panel about Ai

and Drug Discovery thank

you

Get a human-verified training set

Tell us what you're labeling. We'll scope it with our team, yours, or both — and deliver it versioned, in your format.

Also available through Carahsoft and Google Cloud Marketplace.