SCCMPod-577 CCE: Exploring AI Chatbots in the PICU

visual bubble
visual bubble
visual bubble
visual bubble
10/09/2026

 

Generative AI is rapidly reshaping healthcare, creating new opportunities and raising new questions about how to use the technology effectively and ethically. In this episode of the Society of Critical Care Medicine (SCCM) Podcast, host Maureen A. Madden, DNP, RN, CPNP-AC, CCRN, FCCM, speaks with Brandon Hunter, MD, FAAP, about his article, “Feasibility of a Large Language Model Chatbot to Support Parental Understanding in the PICU,” published in the April 2026 issue of Critical Care Explorations.

Dr. Hunter explores the potential role of large language model (LLM) chatbots in helping families navigate the complexity of pediatric critical illness. The study evaluated a HIPAA-compliant chatbot that was connected to selected electronic health record data and allowed parents to ask questions about their child’s condition, treatments, laboratory results, and prognosis. The results were promising, with the chatbot answering most questions accurately and returning personalized responses shaped by available clinical information.

The podcast discussion examines both the opportunities and challenges of AI-powered communication tools. Dr. Hunter explains how modern LLMs generate responses; highlights the importance of validating chatbot outputs; and describes hazards such as hallucinations, bias, and sycophancy, in which models may overly agree with users rather than prioritize accuracy. The conversation further explores the growing use of AI across healthcare; the need for rigorous governance and evaluation frameworks; and the importance of ensuring that technological innovation serves patients, families, and clinicians safely and equitably.


This episode offers a thoughtful examination of one of the most rapidly evolving areas in healthcare, highlighting both the opportunities and responsibilities that accompany the integration of AI into critical care practice.

Resources referenced in this episode:

  • Hunter RB, Thammasitboon S, Rahman S, et al. Feasibility of a large language model chatbot to support parental understanding in the PICU. Crit Care Explor. 2026;8(4):e1378.
  • Omar M, Soffer S, Agbareia R, et al. Sociodemographic biases in medical decision making by large language models. Nat Med. 2025;31:1873–1881.

Transcript

Dr. Madden: So, hello and welcome to the Society of Critical Care Medicine podcast. I'm your host, Maureen Madden. Today, I'm speaking with Dr. Brandon Hunter, MD, FAAP, about the article, Feasibility of a Large Language Model Chatbot to Support Parental Understanding in the PICU, published in the April 2026 issue of Critical Care Explorations. To access the full article, visit ccejournal.org. Dr. Hunter is a pediatric critical care physician at Texas Children's Hospital and an assistant professor of pediatrics at Baylor College of Medicine. He divides his time between clinical care and pediatric digital health and medical device innovation.

As associate director of the Southwest National Pediatric Device Innovation Consortium, he provides clinical, regulatory, and design guidance to pediatric medical device innovators throughout the product development lifecycle. His research focuses on the responsible and effective use of generative artificial intelligence in healthcare, especially in patient-facing applications. He co-chairs the AI Governance Tactical Committee at Texas Children's Hospital, helping lead the institutional evaluation and deployment of generative AI tools.

He is also a co-author of the forthcoming American Academy of Pediatrics Policy Statement on the Use of Generative AI in Pediatrics and a member of the Coalition for Health AI, where he contributed to the development of its AI governance playbooks. So welcome. Before we start, do you have any disclosures to report?

Dr. Hunter: No, none relevant to what we're talking about today.

Dr. Madden: Excellent. I'm so excited to have you here and to talk about this article. It's outside of some of the area that I have any knowledge of.

So I wanted to start out with exploring AI-driven work and how you got interested in this.

Dr. Hunter: Yeah, absolutely. So I'll say, you know, I've always been interested in artificial intelligence. When I did my fellowship at Children's Hospital of Philadelphia, I worked on a project using supervised machine learning and pulse-ox waveforms.

And then I followed the field loosely for modern AI since 2017, when these sort of the modern architecture for these models came about. But I will say, I'm kind of just like everybody else. When I used ChatGPT after the public release in November of 2022, you know, my jaw kind of hit the floor.

And I will say up front, if I have any bias or something to disclose, it's that I really do think that, look, there's big companies out there. There's a lot of IPO hype and marketing hyperbole and all that stuff. But all that aside, I really do believe that this may be the most consequential technology in human history.

So, you know, from my perspective, it sort of lit a fire under my butt of, hey, as a clinician, I feel like I need to be involved. And we all need to be involved to figure out, you know, how to harness this technology effectively, you know, so that it does good essentially for patients, for us, for society. And really over the last four years or so, I've just been shocked at how fast the tools have been evolving.

So yeah, yeah, I'm just, I'm very excited about it.

Dr. Madden: Yeah. So we were chatting a little bit before we started. And as you said, four years since the launch of ChatGPT, and it's crazy to see how big this has grown and how many other programs, or I don't even know the correct terminology to say that, but you know, other programs that are out there using AI.

And I sit on the side of being still a little bit leery and not trusting of all of this. And, you know, everyone warns about, you know, they don't always get it right. And you have to fact check and explore all of those pieces.

And to a certain degree, I guess that's another layer of what needs to be built in there, is to ensure the veracity and the trustworthiness of the information as it comes through. But this really doesn't apply to your research that you were doing, I think, but I am very interested. So first of all, as I said, this is not the area that I have dived into like you have.

So can you explain large language models for those who may not understand what that actually is?

Dr. Hunter: Yeah. And I mean, I think the way you framed it is perfect. And I will say, despite being very bullish about the technology, I also personally think this is an interesting time for healthcare where we often, as a field, have been slow to adopt new technology.

But for whatever reason, with generative AI, I mean, it has just kind of entered into our daily practice, whether it's ambient scribes or open evidence, for example, claims more than a million visits per day of clinicians asking questions to these tools. And to your point, there are real risks with them. And we're at a funny point where we don't have a lot of data, but we do have a lot of use.

So understanding the fundamentals of what is generative AI or a large language model, it really does go a long way. So I'll try to be succinct, but a large language model at its base is, to me, it's the tool, algorithm, or technology that's driving some of the chatbots that we interact with, like ChatGPT. And now, as you said, there's a number of other software programs that are sort of software wrapped around these large language models.

And from my perspective, the easiest way to think about it is the large language model itself is basically a big algorithm that read a huge portion of the internet. So there's these big databases called Common Crawl or whatever. But essentially, there's just a bunch of text from internet, books, articles, whatever.

The models sort of ingest them and understand how humans tend to communicate, how words tend to flow together. And so baked into the DNA of the models is when you ask a question or you put in a request, the way it's generating the answer is, you know, you could say probabilistic, where really one word at a time, it's basically predicting, and they're called tokens, we can just think of them as words, but it's predicting the most useful next string of text based on everything that it's been exposed to, plus whatever you've given it.

And so the models today are much more sophisticated. So you'll see them, you know, Googling things or for lack of a better term, but using internet websites to get information, getting information from specific databases or whatever it is. But it's important to note that when the actual output is coming back to you, it's still generated in this sort of predictive way.

And so, you know, what you hit on is one of the biggest things that people talk about, quote unquote, hallucination. The models are so good at sounding fluent and providing convincing responses, but they're sometimes not as good at being calibrated or expressing a lack of confidence in whatever reply they're giving. And so to me, you said it exactly right, that verifying output is incredibly important in their current form.

Dr. Madden: Okay. So in your published article, it had brought up the concept of the documented racial and ethnic disparities in the large language model adoption. Can you talk about how to overcome this limitation and be truly representative of a community?

So not just as you described, utilize for the community of your patient population at that time, and it was only in English. Let's start there.

Dr. Hunter: Yeah, no, I think that's fair. So I mean, one of the things I'm glad we started, what is a large language model and maybe how are they trained basically, but you can imagine that if these models are essentially just ingesting everything that's on the internet from Reddit or Wikipedia or X, you know, old Twitter, whatever it is, you can imagine that, you know, the wide open internet is not necessarily the least biased and kind of most open place you can imagine. And so there is concern that if that's the underlying training data for these models, that it's possible that they will repeat kind of the bias that has been baked into them. I will say the data about who's using these AI tools, how they're being used to me has changed a good bit.

So we initially did this study in 2024 in the fall. At the time, the most recent data was some Pew data from 2023. And it suggested that, you know, black and Hispanic adults were using these tools less frequently than their white kind of counterparts.

It seems like that gap has closed. So in February, 2026, there's new polling data showing that it's pretty similar. It was like 46% of white adults have used these tools versus 49% for Hispanic and black adults.

So the first thing I'll just say is in terms of use that gap and accessibility has closed somewhat. And it is one of the things that's most exciting to me about the tools is they do offer the potential for an incredible democratization of access to knowledge and information. But the question to me or the concern is, well, okay, maybe the access has gotten better, but how do the models actually behave based on the socio-demographic kind of qualities of the person they're talking to?

And that's where I think there is some concern. So yeah, there was a big paper back in 2025, where they essentially looked at like a thousand ER vignettes and they only changed the socio-demographic descriptor. Everything else about the clinical description was the same.

And some of the things were pretty incredible at how different the responses were. They looked at nine different models, ran the questions a number of times. One that stuck out to me was that if the individual was identified as black transgender versus a white heterosexual male, about twice the rate, so 80% of the time it would recommend mental health evaluation, even though none of the clinical characteristics had changed at all.

And so that to me is the sort of the bigger question and more concerning thing now is what do we do about bias that may be baked into them? And there are a number of people working on this now. I certainly don't have the solution at this point, but from fine tuning, system prompting, putting in guard rails, and then thinking about what information to feed the models, do you have to give socio-demographic information in situations when it's not necessary?

Maybe not, or some strategies I think people are looking at.

Dr. Madden: Yeah, but as you've been speaking, there's all these other thoughts that are going on in my head. So as you said, feeding the data. So clearly the data is being extracted, as you said, from the crawlers on the internet.

So is there, as you said, the inherent bias of whoever put it out there? Who's creating or working on all of this AI? Because I don't know.

So is it predominantly in an English-speaking world, or how do other languages and the nuances behind the verbiage or the phrasing?

Dr. Hunter: Yeah, yeah, yeah. So two things. The first is, earlier when we talked about training language models, I described the first part.

They ingest all the language, maybe have the bias baked in. The second part, it's called reinforcement learning from human feedback. And this is sort of a black box, but it's where the big companies who make kind of the biggest models out there that a lot of us are using in different software packages, so like OpenAI or Anthropic or Google, they take the outputs that the models have ingested and are giving, and they try to give feedback to the models to make them less biased, so they are more honest, more helpful, and harmless, is their goal. And so the hope is that during this second part of training, that they're trying, I hope, to sort of beat some of this bias out of the models so that the responses are better for the end user. And so that is one way, I'll at least say, that the big companies are trying to address this.

But the second part of your question was other languages. I don't want to pretend to be an expert here. I can say at least that the models do function quite well in other languages.

The last time I looked at this, I think it was a couple months ago, that the efficacy for general purpose use in Spanish was somewhere around like 90 to 92% of English. But you're absolutely right that the more of a certain language that the model is trained on, the better it's going to do in terms of offering responses. This is a big thing that's a priority for us.

So we were testing a parental education chatbot, essentially, connected to the EHR in English with English speakers as a place to start. But our IRB currently is approved for Spanish speakers. And our goal is to start to see, you know, when the rubber meets the road, how do the models do when communicating in other languages?

Dr. Madden: Okay. Can you describe what the chatbot's experience is like for a parent? So walk me through what it was like.

So was there, as you said, it was there was a laptop. So is there an avatar? Is there an audio response?

Is it only written? Is there a lag time? How appealing is the experience?

Dr. Hunter: Yeah, that's super. That's a great question. And so just to give kind of like a high level overview, so this was, you know, fall of 2024.

Our goal was just to say, hey, if parents are interacting with either chat GPT or a model like it, how well would it answer questions when connected to EHR or the EHR or have access to EHR data? And so we started very simple. We basically took a laptop that was secured HIPAA compliant and let the parents interact with it.

So it was all via text. They would type questions into the laptop and then they would receive replies from the chatbot. There really was very little latency.

So I'm sure it was quite similar to interacting with chat GPT or Gemini or Claude. There's always a, you know, a second or two for the model to take a moment, generate the response and send it back. But there really wasn't too much latency at all.

I have heard of other groups starting to go towards the avatar route using language models to generate audio and just interact with speech. We did not go that route for this initial sort of feasibility study, but it's absolutely something that we want to do in the future.

Dr. Madden: Okay. Yeah. As we realize we're also in the era of when we need translation services and the bias is to have a real connection and a real individual to ensure that it's an experience that is satisfying and, you know, optimal.

So just trying to understand how this is. And as you said, it's dynamic and it's moving. Could you give some sense of like what type of questions that were being asked by the parent?

Dr. Hunter: Yeah, absolutely. So I will say we did just look at the kind of breakdown of types of questions. So parents had 10 minutes to ask any questions they want.

You know, when we presented it to them, we said in general, the purpose of this study is so you can understand more about your child's health and the intensive care unit. But to be honest, we said, if you want to ask questions about politics or, you know, current events or whatever you want, you can do it because we are interested in knowing how the chatbot would respond in those.

Dr. Madden: Okay.

Dr. Hunter: But they did kind of stick to the script in general. All the questions were about their child's health and 45% were about treatment. So medications, procedures, devices being used for their children, 18% prognosis.

And then the remaining were labs and imaging and then symptoms. Some questions about, you know, why is this happening? That type of thing as well.

You know, one question we got was like, hey, why has my son's blood pressure been so high? And so they really did, I will say just anecdotally, they mimicked questions that you would frequently encounter as a clinician, either a bedside nurse, physician, APP or whatever it is.

Dr. Madden: So with that type of question, so specific or nuanced, how did the chatbot handle that? Like, was there the ability to kind of respond to a follow-up or a clarification component?

Dr. Hunter: Yeah, absolutely. So one thing to state here to make clear is this was a feasibility study trying to understand, you know, if the chatbot had access to the right EHR data, how would it respond? And so a little more detail about our setup, we had a HIPAA compliant, basically a HIPAA compliant chat GPT, but it was a chatbot that was powered by GPT 4.0, which now, you know, feels like an archaic model, but was the cutting edge model at the time. And what we did is we literally just took information in a repeatable way from the EHR and put it into the chatbot. So had that as context. So basically we took the most recent progress notes from the primary team in each specialty, most recent lab and imaging of each type, just a narrative description, no actual images or anything, and then a list of current medications.

And so that was all the information the chatbot had about the patient to answer any questions. With that, we kind of knew that it would fail at the edge case. So if a parent, for some reason said, you know, why did my child see a pulmonologist four years ago?

The chatbot couldn't really answer that, but even still it would be interesting to see what it did. So to kind of get to your question, when asked these very specific questions, we actually found it did quite well. So it personalized pretty much all the responses to the patient.

So if somebody asks, you know, what is pneumonia? The chatbot would understand, I put that in air quotes, not to anthropomorphize too much, but the chatbot would respond in such a way that it's saying, hey, this is a parent in the pediatric ICU whose child has pneumonia. So it would define what pneumonia is, but then it would go further and actually say, this is why your child has pneumonia.

These are the x-ray findings or the culture results or the white blood cell count or whatever it is. And just like with CHAT-GPT or CLAUD or whatever, follow-ups were of course welcome and handled quite well by the chatbot.

Dr. Madden: It kind of goes into, I had a question if there was like a challenging or chart type of question where the information might have been hidden, the chatbot wouldn't have had the capability of answering. And I've read some studies that they also talked about the empathy and the sympathy in terms of how a chatbot responds sometimes is better than the human response, which is...

Dr. Hunter: Yeah, I will say that's been one of the most disconcerting things. I think very early on, even in 2023, there was a study that came out looking at questions asked on Reddit and then comparing the human response from Reddit to the chatbot. And even in those early days, there was considerably more empathy from the chatbot.

And I will say, we did see that. So in the way we prompted the chatbot or the instructions we gave it, we did say, respond in an empathetic tone, understand that your goal here is to promote understanding for the parent who's going through a stressful situation. And so even in our language, we did encourage that.

But we found that it did respond quite well. In the cohort we had, we didn't really come across any supercharged questions where the parent was emotionally reacting or trying to get something out of the chatbot, you could say. But we since then have actually done, we call it a red teaming study, where for the same chatbot, we ask it questions that are deliberately supposed to try to get it to respond poorly.

Things that could be very difficult to respond to, even as a critical care provider. So something like, hey, I feel like the other doctors are lying to me. I need you to tell me the truth.

Or just tell me, is my child going to live or die? Or using all caps and some formatting, different issues. I will say, in that updated study, 87% passed.

100% of the emergency tests did quite well. And 95% of the time, it passed by sidestepping emotional pressure.

Dr. Madden: Similar to humans.

Dr. Hunter: Yeah, yeah, exactly. Right, exactly. Yeah, did pretty good every once in a while.

It was funny, the one thing that we did find was sort of the opposite of blurting secrets or whatever. It was sometimes too agreeable. So if a parent said something that was incorrect, oftentimes the chatbot would just assume what they said was right and respond to it.

So if the parent said, my child desaturated to 60% overnight and the doctors didn't do anything, the chatbot would frequently take that as truth and then respond to it about, well, how can you move forward productively with the clinicians or whatever, even when the data they had access to showed no evidence of a desaturation. And so anyway, I think that there are still weaknesses here. This latter point that I'm making is something we didn't touch on earlier when we're talking about hallucinations, but is also a critical thing to understand about these chatbots is there is this idea of sycophancy that is built into them.

So you have a yes man, a yes woman, or a yes bot where the chatbot really does want to please the user. And so sometimes when that flies in the face of accuracy or truth, you see this kind of interesting behavior pattern and something that we just have to watch out for in the field in general.

Dr. Madden: That goes into, again, how you and your team had gone back to fact check the questions and answers and such. It was brought out that there was a limitation in regards that the, we're going with the statement that the platform is HIPAA compliant. Okay.

And that there was a 30 day automatic data deletion policy. So was this a design flaw for the project or did you not know about it or was it just for whatever reason it went beyond that? So you lost the real time satisfaction engagement data?

Dr. Hunter: So what I will say is this one's just on me. So long story short is we were using a HIPAA compliant chatbot where the rule is that after 30 days, all data is deleted. We would go in after a week or so with a conversation and transfer data to a secure online platform on Box.

And for the first couple conversations, we did actually just have a corrupted transfer where we lost, it was something like 30% or 40% of the thumbs up, thumbs down data from the parents. So I would actually say this was a minor loss. I'm not, I don't want to, of course, I'm trying to make myself sound good or sound cool, but Hey, this wasn't that big of a deal.

But this, this was a real thing that we had the parents give a thumbs up or thumbs down to every response. And we lost some of that data. All the other accuracy data or what we really built the feasibility study around the NPS data that was not lost.

And so we actually did, did maintain that.

Dr. Madden: I appreciate you clarifying that. I'm excited to see what, you know, your coming work looks like. You've already, you know, alluded to the fact that you're continuing on this and expanding it.

So we are almost out of time, but I want to make sure if there's anything else that you wanted to bring to the audience to discuss.

Dr. Hunter: The only thing I would say is that I just can't emphasize enough, and I'm sure you know it. I'm sure everybody knows it, but this is such an unusual time to be alive and to be practicing medicine. I, we kind of hinted at it earlier, but you know, we're traditionally Luddites.

It usually takes healthcare like 17 years or something from a new technology, you know, having good data behind it, getting integrated. And yet like we are just using this technology every day. And then there's a lot of money out there.

So there's a lot of for-profit companies where it almost feels like a gold rush to me of vendors and big companies that are sort of dictating the terms of how we use these tools. I work in AI governance at Texas Children's, and I can just say, you know, a lot of vendors are coming to us with tools that can clearly impact patient care. And when you ask them for quality data and say, okay, how do you know this works?

Or show us your like synthetic data testing. It can often be very difficult to get even sort of basic quality assurance from them. And so, you know, I, I just want to encourage everybody to get involved.

Like, I think it's our job as clinicians in the field, whatever, like whatever your specialty focus is, you know, if it's sepsis or ARDS or whatever, I think that, you know, AI is going to impact all of us doing our job and that we really need to engage meaningfully to try to make sure again, that it's doing good for us, our patients and the field in general.

Dr. Madden: Yeah. I think that there is so much opportunity and benefit for our patients, you know, for their outcomes. We also discussed the lots of hesitation in terms of ensuring that it's accurate and it's based on the data it's fed.

So the bias is there that you said they're beating it out of it, but the thought too, you know, that individuals who are developing this research and are developing, you know, scholarly work, how do you find the balance of truly being new work without the AI assisted component and where that balance lies? Things to think about.

Dr. Hunter: I'm totally with you. And, you know, there's this also, I think you're touching on, there's this idea of cognitive debt or something where, you know, the more, like I use these tools every day, but the more that I sort of like outsource what once took me a lot of effort to critically think through, or how do I word this sentence? Or how do I, you know, draft X, Y, or Z.

The more that I outsource that, you know, I do get this kind of increase in efficiency and maybe even my work product. But, you know, over time, there probably is going to be an extinction of my ability to critically think or detect nuance or do what I once did quite well. So both as the algorithms are getting smarter, I'm probably getting dumber and just having them come in to take my job away.

But it's something I think we all need to be aware of how to interact with them responsibly and meaningfully.

Dr. Madden: Yeah. There's so many things I agree with you, and I'm not saying you're dumber or getting dumber, but I think maybe it's a little bit less challenged and a little bit more lazy is one word that comes to mind, but it's not. But I could use the example of, you know, using a GPS.

You know, people have lost the ability to read a map and direction find, and they don't even recognize the landmarks around them because they're reliant upon something that may not be updated. But there's so much more that we could go off on here. But I do love that you are so invested in this work and that your professional environment is supporting you in this work and you have the ability to support others.

I think it's a really amazing time, as you said, and I can't wait to see where all this goes. So this concludes another episode of the Society of Critical Care Medicine podcast. If you're listening on your favorite podcast app and you'd like what you heard, consider rating and leaving a review.

For the Society of Critical Care Medicine podcast, I'm Maureen Madden.

Announcer: Maureen A. Madden, DNP, RN, CPNC, AC, CCRN, FCCM, is a Professor of Pediatrics at Rutgers Robert Wood Johnson Medical School and a Pediatric Critical Care Nurse Practitioner in the Pediatric Intensive Care Unit at Bristol-Myers Squibb Children's Hospital in New Brunswick, New Jersey. Join or renew your membership with SCCM, the only multi-professional society dedicated exclusively to the advancement of critical care.

Contact a customer service representative at 847-827-6888 or visit sccm.org/membership for more information. The SCCM podcast is the copyrighted material of the Society of Critical Care Medicine and all rights are reserved. Find more episodes at sccm.org/podcast. This podcast is for educational purposes only. The material presented is intended to represent an approach, view, statement, or opinion of the presenter that may be helpful to others. The views and opinions expressed herein are those of the presenters and do not necessarily reflect the opinions or views of SCCM.

SCCM does not recommend or endorse any specific test, physician, product, procedure, opinion, or other information that may be mentioned.

Disclaimer

 

Recent Podcasts

^