
Eine neue Studie ergab, dass LLMs wie ChatGPT in 50 bis 82% der Fälle auf gefälschte klinische Details verlief. Selbst die besten Eingabeaufforderungen konnten nicht alle Halluzinationen aufhalten. Der klinische Gebrauch benötigt immer noch starke Schutzmaßnahmen.
https://www.nature.com/articles/s43856-025-01021-3
10 Kommentare
Well yeah, it’s designed to respond to prompts, no matter what the prompt is it’s going to make something up. People expecting LLMs to be like calculators is intriguing.
Is anyone using LLMs for clinical use?
[deleted]
Maybe it asked Google search engine?
Meanwhile, the top news story on Australia’s ABC news site:
https://www.abc.net.au/news/2025-08-06/doctor-ai-artificial-intelligence-scribe-notes-appointment/105615244
Clinical use needs safeguards? No. LLMs **shouldn’t be used in a clinical setting** until they work well enough. And since these models are training on their own output, who knows if they ever will.
I think it’s equal to humans
*Large language models (LLM), such as ChatGPT, are artificial intelligence-based computer programs that generate text based on information they are provided to train from. We test six large language models with 300 pieces of text similar to those written by doctors as clinical notes, but containing a single fake lab value, sign, or disease. We find that the LLM models repeat or elaborate on the planted error in up to 83 % of cases. Adopting strategies to prevent the impact of inappropriate instructions can half the rate but does not eliminate the risk of errors remaining. Our results highlight that caution should be taken when using LLM to interpret clinical notes.*
ChatGPT like LLMs are designed to tell you the answer you want to hear, not the truth.
It is pretty simple: if you want it to detect fabricated input, you must train it to detect fabricated input.