Every published Top AI Stories item tagged with Health Care & Social Assistance, newest first.
The I3LUNG project, a study across six centers in Italy, Germany, Greece, Israel, Spain and the United States, reported in Nature Medicine that AI models built from 2,396 patients with advanced non-small cell lung cancer predicted immunotherapy outcomes better than every standard clinical biomarker, including the PD-L1 test doctors rely on. When 20 physicians reviewed 100 real cases with the tool, their accuracy in spotting responders rose substantially, with the largest gains among doctors who were not lung cancer specialists. This phase looked back at past patients; a prospective trial of more than 2,000 is enrolling.
Google DeepMind released AlphaGenome Atlas, running its AlphaGenome model across all nine billion possible single-letter changes in the human genome and publishing the predictions as a one-petabyte dataset, more than 30 times the size of the AlphaFold database. It is free for non-commercial use, with commercial access on Google Cloud to follow. Early collaborators, including the Broad Institute, Boston Children's Hospital and Memorial Sloan Kettering, report finding 22 percent more non-coding genetic associations and pinpointing a variant in the DNM1 gene tied to a severe childhood epilepsy.
OpenAI opened an Epic integration that pulls authorized patient context — visit notes, laboratory results, medications and specialist documentation — into ChatGPT for Healthcare, so a clinician can ask what changed since a patient's last visit instead of reading through the chart. A companion plugin adds structured access to nine official public sources including PubMed, ClinicalTrials.gov, DailyMed and Medicare coverage policy. The University of California, San Francisco health system is a pilot partner. OpenAI says physicians across 60 countries have reviewed more than 700,000 model responses to tune its healthcare behavior.
Researchers at Chaim Sheba Medical Center and Tel Aviv University trained a deep learning model on 97,364 mammography examinations from 29,921 women, median age 54, and found it could pick out cardiovascular disease from the same images taken to screen for breast cancer. Measured as area under the curve, detection reached 0.86 for prior stroke, 0.79 for high blood pressure and 0.78 for ischemic heart disease. This is a research prototype rather than a cleared device — the work is a conference abstract due to be presented at the European Society of Cardiology meeting in Munich, and the team says it is still working to cut both false positives and false negatives.
Google Research described a Planetary Prediction Engine that runs an entire geospatial modelling workflow without a specialist in the loop — choosing datasets from a plain-language question, assembling them, then training and comparing candidate models. Applied to an Ebola outbreak in the Democratic Republic of the Congo, it correctly picked out 15 of the 18 health zones the virus newly reached, about ten percentage points better than the statistical method it was measured against. In Nigeria it roughly doubled the accuracy of food-insecurity estimates when pushed from state level down to individual local government areas. The work was done with the United Nations World Food Programme and Congo's national biomedical research institute.
A peer-reviewed study in PLOS Digital Health traced every AI-enabled medical device the US Food and Drug Administration had cleared through December 2025. Of 1,357 devices, 34 were linked to a registered clinical trial, 12 reached a peer-reviewed publication, and just three measured whether patients were actually better off — mortality, illness, or readmission. Roughly four in five American hospitals now use AI somewhere in care, so the distance between what clears review and what anyone has shown works is wide and growing.
Google Research and MIT built a multi-agent system that generates hypotheses, runs the statistics and checks the literature to surface candidate biomarkers hidden in wearable sensor data. Across three cohorts totaling 9,279 participant-observations it proposed 41 candidates for mental health and 25 for metabolic disease, including a link between night-to-night sleep variability and depression severity. In a blinded review, domain experts scored it highest on all seven quality measures. These are leads for further study, not a clinical tool.
Google's PhotoScan model estimates body composition — including the visceral fat ratios that a scale and a tape measure cannot see — from standard two-dimensional smartphone pictures. It was pre-trained on more than 35,000 UK Biobank participants with clinical scan ground truth, fine-tuned on 677 adults photographed with real phones, and validated against a 132-person longitudinal trial. It beat smartwatch bioimpedance on body fat percentage and improved insulin-resistance classification well past demographics alone. Google calls it a research prototype, not a product.
Google Research reported that AMIE (Video), a version of its medical dialogue system that handles live audio and video, was rated on par with board-certified primary care physicians on history-taking, diagnostic accuracy, management and communication. The randomized study covered 100 clinical scenarios and 300 consultations with trained patient actors, scored by 20 independent physicians. AMIE rated significantly higher than the doctors at eliciting physical signs and guiding examinations. It remains research-only — every consultation used actors, not real patients.
A Stanford and Arc Institute team led by Brian Hie used the Evo 1 and Evo 2 genome language models to write complete viral genomes from scratch, then built about 300 of them in the lab. Sixteen produced working bacteriophages — viruses that infect bacteria — and three beat the natural template they were modeled on, while a cocktail of the designs broke through E. coli strains that had already evolved resistance. The models were trained with human viruses excluded from the data, and the authors argue that exclusion is what makes the work publishable; several reviewers responded that the harder question is what happens when a lab without that restraint repeats it.
When Fable 5 shipped, its biology classifier was tuned so cautiously that it pushed a wide range of harmless questions down to Claude Opus 5, and Anthropic accepted that cost because a frontier biology model could give real uplift to someone building a weapon. The company has now rewritten the classifier's constitution, carved out explicit exceptions for benign cases, retrained on new data, and reports about 85 percent fewer biology-related fallbacks across its products. Virology, toxicology and molecular design still get routed away — the loosening covers reading lab results and learning biology, not dual-use research.
Researchers at MIT led by Marzyeh Ghassemi, with collaborators at Stanford and Columbia, tested skin-disease diagnosis with and without model assistance. Non-experts did get more accurate, but largely by deferring: they trusted the model's written explanations whether those explanations were right or wrong, and rated the vaguest ones as most convincing. Clinicians were not misled by incorrect assistance and performed best when handed a bare prediction with no explanation at all — which inverts the usual assumption that showing the reasoning makes medical AI safer.
Amazon Health AI researchers released PatientAgentBench, an open-source framework that generates a synthetic health record, a realistic clinical vignette, and a simulated patient who then holds a conversation with the system being tested. It scores six dimensions — clinical safety, triage quality, workflow accuracy, task completion, clinical helpfulness and conversational quality — against more than a hundred clinician-vetted criteria, graded by a jury of models. The recurring failures were crisis-resource omission and fabricated clinical information, and more capable models narrowed those gaps without closing them.
Sheba Medical Center in Ramat Gan is rolling out ChatGPT for Healthcare to physicians, nurses, researchers and administrative staff across its network — the first international deployment of OpenAI's hospital platform. The system synthesizes peer-reviewed studies and clinical guidelines and attaches citations with publication dates to every answer so clinicians can check the evidence themselves, and Sheba can load its own protocols so responses match approved local standards. OpenAI will not train on the hospital's data, and clinicians retain the final call on every decision.
Researchers used an AI protein-design model to invent synthetic versions of a compact gene-editing enzyme — a small cousin of the CRISPR proteins — and the best designs beat the natural reference. In human cells, two AI-designed variants hit 46 and 50 percent editing efficiency versus 28 percent for the original, with nearly fourfold gains at some targets. Published in the journal *Science*, the work shows generative models can expand the gene-editing toolbox with smaller, more effective enzymes — a boost for future gene therapies where delivery space is tight.
OpenAI made ChatGPT Health available to every US adult across its free and paid tiers, letting users connect Apple Health data and medical records from systems like Epic and Oracle Health for personalized guidance. The company says more than 300 million people now ask ChatGPT health questions each week, up from 230 million in January. The wider rollout lands a day after a Florida pastor sued OpenAI, alleging the chatbot gave a near-fatal suggestion to skip a doctor — a reminder of the stakes as general-purpose models move into medical advice.
Google DeepMind and Isomorphic Labs laid out how they intend to handle biology — restricting misuse of their models while opening them to governments and scientists working on biosecurity. The two describe more than 15 partnerships with government bodies, biosecurity organizations, and research groups, and point to work already running: AlphaFold and AlphaGenome for pathogen characterization, AlphaEvolve tuned for metagenomic sequencing to catch outbreaks earlier, and a unit inside Isomorphic aimed at designing countermeasures quickly when a novel pathogen appears. Extending SynthID-style watermarking to screen DNA synthesis orders comes next.
Engineers and surgeons at UC San Diego used teleoperated humanoid robots to complete two minimally invasive gallbladder removals on live pigs — a world first reported in the journal Nature. One operation paired a robot with a human surgeon acting as assistant; the other was run by two robots working side by side, both wielding standard surgical tools rather than custom hardware. If the approach reaches human patients, it could bring robotic surgery to smaller hospitals that cannot afford today's specialized systems.
Google Research introduced SensorFM, a foundation model pre-trained on more than one trillion minutes of sensor data from five million people wearing Fitbit and Pixel Watch devices. It draws on signals like heart rate, blood-oxygen, sleep, motion, and skin temperature, and transfers to 35 health-prediction tasks across cardiovascular, metabolic, sleep, and mental-health domains. When it grounded a personal health agent, its predictions matched the quality of using the real measurements.
An IEEE Spectrum report tracks how small models running directly on phones and cheap hardware are delivering results in regions with no reliable internet or data centers. The anchor: RxScanner, founded by Adebayo Alonge, pairs a handheld spectrometer with a phone-based AI model to spot counterfeit pills — a problem that kills thousands each year — and now runs entirely on Android after cloud latency proved unworkable, with deployments in Ghana, Kenya, Myanmar, and Nigeria. World Bank president Ajay Banga is backing the approach with grants and policy support, aiming to close a gap where only 0.7 percent of internet users in the poorest countries have tried tools like ChatGPT versus 25 percent in wealthy ones.
Meta and Spain's Basque Center on Cognition, Brain and Language published Brain2Qwerty, a deep-learning system that decodes typed sentences from non-invasive brain scans — reaching 61 percent word accuracy, and 78 percent for its best participant, versus roughly 8 percent for earlier non-invasive methods. It still needs a room-sized brain scanner, so it is a research milestone rather than a product, but it points toward communication tools for people with brain injuries that skip surgical implants. Meta released the code and dataset.
xCures closed a $46 million Series B led by Innovius Capital to expand its "Clinical Clarity Engine," which uses AI to turn scattered, unstructured medical records into decision-ready data for doctors. Spun out of Cancer Commons in 2018 to help advanced-cancer patients find treatments, the company says it has now processed more than 300 million records from over 550,000 sites across every clinical field. Patient histories are often trapped in mismatched documents across labs, hospitals, and imaging centers; xCures aims to make them usable at the point of care.
Immunologist Derya Unutmaz at The Jackson Laboratory handed GPT-5 Pro an unpublished dataset his team had puzzled over since 2022 — how glucose shapes the way T cells develop and specialize. Within minutes the model pinpointed the likely mechanism, disrupted N-linked glycosylation, and proposed follow-up experiments that later confirmed it. OpenAI published the case on June 23 as evidence that frontier models can now generate original, testable scientific hypotheses, not just summarize what is already known.
Assort Health raised a $120 million Series C led by Menlo Ventures at a $1.2 billion valuation, becoming a unicorn on the strength of AI voice agents that run the patient journey — scheduling, intake, referrals, medication refills, and payments. The company says it has now handled more than 190 million patient interactions and grown revenue 20-fold over the past 15 months. The round is a marker of how quickly health systems are handing routine front-desk work to AI rather than hiring for it.
The AI image company Midjourney unveiled a medical division and its first hardware: an "Ultrasonic CT" scanner that uses 500,000 transducers and no radiation to image the whole body in about 60 seconds, reconstructing the scan with AI. Midjourney licensed Butterfly Network's ultrasound-on-chip technology and plans to debut the device in a San Francisco "spa" in late 2027, aiming for 50,000 machines worldwide by 2031.
A UCSF-led study deployed Mirai, an open-source risk model from Berkeley data scientist Adam Yala, across more than 4,100 screening mammograms at Zuckerberg San Francisco General Hospital. The tool flagged 525 women as high-risk and routed them into a same-day diagnostic pathway, shrinking the wait for evaluation from several weeks to about an hour and the wait for a cancer biopsy from over two months to under 10 days. Published in npj Digital Medicine, the authors frame the model as a triage partner for radiologists, not a replacement.
GE HealthCare won FDA clearance for MIM Contour ProtégéAI+ 2.0, software that automatically outlines tumors and healthy organs on CT and MR scans — one of the most time-consuming steps in planning radiation therapy. By handling the contouring with minimal input, it frees oncology teams to spend more time tailoring each patient's treatment. The clearance also includes a change-control plan that lets GE push future updates to new body regions faster, a sign regulators are adapting to clinical AI that keeps improving after approval.
University of Cambridge researchers began human trials of what they call the first vaccine whose central component was designed entirely by AI. The shot targets the whole Sarbeco coronavirus family — including Covid-19, SARS, and related bat viruses — in hopes of guarding against outbreaks that have not yet emerged. An early 39-person trial tested safety; a roughly 200-person study will gauge immune response. The team is already applying the same AI design approach to influenza and Ebola vaccines.
OpenAI announced its Rosalind Biodefense Program on Friday, opening GPT-Rosalind — a frontier model fine-tuned for life-sciences work — to select US government agencies, allied nations, and a vetted set of private developers. Named applications span epidemiological modeling, early detection, screening, preparedness, non-pharmaceutical interventions, and medical countermeasure development. OpenAI says it briefed the White House and federal public-health agencies on its approach, though it has not yet published red-team or misuse-prevention details for the dual-use model.
Anthropic announced a four-year, $200 million commitment to the Gates Foundation, structured as a mix of cash grants, Claude usage credits, and technical support from its Beneficial Deployments team. Focus areas span global health (polio, HPV, and preeclampsia and eclampsia among named disease targets), life sciences, education, and economic mobility, with regional emphasis on sub-Saharan Africa, India, and other low- and middle-income countries. Launch partners include the Institute for Disease Modeling and the Global AI for Learning Alliance. It is the largest single philanthropic commitment from a frontier lab to date and a meaningful structural signal that Anthropic intends to be measured by social-impact deployments alongside its commercial book.
CMS launched ACCESS (Advancing Chronic Care with Effective, Scalable Solutions), a 10-year initiative starting July 5 with 150 participating organizations and outcome-based payments that, for the first time, reimburse organizations for AI agents monitoring patients between visits, coordinating referrals, and managing medication adherence across six conditions including diabetes, hypertension, and depression. The catch: reimbursement rates are low enough that "the math only works for organizations that have fully automated most patient interactions," per Pair Team's CEO — making AI not just allowed but operationally required for participating providers.
Pennsylvania filed suit against Character.AI after state investigators interacted with a chatbot named "Emilie" that claimed to be a licensed psychiatrist and produced a fabricated state medical license number while discussing depression treatment. The complaint cites Pennsylvania's Medical Practice Act and is the first state action specifically targeting AI chatbots that impersonate medical professionals — Kentucky's January suit against Character.AI focused on harm to minors, not credentialing fraud. Expect more state attorneys general to follow this template.
On May 4 both frontier labs announced finance-industry-backed enterprise AI services companies. Anthropic revealed a joint venture with Blackstone, Hellman & Friedman, and Goldman Sachs — also backed by Apollo Global Management, Sequoia, GIC, General Atlantic, and Leonard Green — that embeds Applied AI engineers into mid-sized community banks, manufacturers, and regional health systems for Claude integrations. OpenAI's parallel "The Development Company" raised $4 billion at a $10 billion valuation alongside TPG, Brookfield, Advent, and Bain Capital. Both adopt Palantir's forward-deployed engineer model — sending lab engineers into client organizations rather than selling SaaS — a structural signal about how mid-market AI distribution is going to be sold.
A peer-reviewed paper in *Science* pitted OpenAI's o1 reasoning model against two attending physicians on 76 emergency-room triage cases. The model landed on the exact or near-exact diagnosis in 67 percent of cases, versus 55 percent and 50 percent for the two doctors; on a wider 143-case cohort the correct answer sat inside o1's differential 78 percent of the time. The authors caution the inputs were text-only EHR snippets and call for prospective real-world trials before any clinical deployment.
Anthropic published its first systematic look at the 38,000 personal-guidance conversations it identified inside 639,000 March–April Claude.ai chats. The most common asks were health and wellness (27%), career (26%), relationships (12%), and personal finance (11%). Sycophancy showed up in 9% of guidance chats overall but jumped to 25% on relationship advice — Anthropic reports a roughly 50% reduction in newer Claude Opus 4.7 and Mythos Preview training runs.
DeepMind announced an "AI co-clinician" research initiative built on Gemini and Project Astra, structured as a triadic care model where the AI works alongside a supervising physician with a separate "Planner" module monitoring safety boundaries. Academic collaborators include Harvard Medical School and Stanford Medicine, with phased trusted-tester evaluations planned across the US, India, Australia, New Zealand, Singapore, and UAE. It is research, not yet approved for clinical diagnosis or treatment.