A machine learning model that combines liver and spleen stiffness measurements with routine clinical data identified clinically significant portal hypertension (CSPH) in patients with compensated advanced chronic liver disease without invasive testing, leaving fewer patients with indeterminate results than current criteria, according to a study published in the Journal of Hepatology. In an external validation cohort, inconclusive results fell to 12% from 48%.
“These findings are encouraging, but they require dedicated prospective validation before the model can be considered a prognostic tool,” said one of the study authors, Antonio Colecchia, MD, professor of gastroenterology at the University of Modena, Italy, who spoke with GI & Hepatology News. “Overall, this work represents a meaningful advance in non-invasive hepatology by harmonizing elastography measurements and integrating them with key clinical and biochemical variables.”
The study validated the pan-elastography machine learning score, known as the ELM score, in 1,435 patients from 24 international centers. All patients underwent hepatic venous pressure gradient (HVPG) measurement, liver and spleen stiffness testing, and routine clinical evaluation. CSPH was defined as an HVPG of at least 10 mm Hg. The investigators compared the ELM score with the current Baveno VII criteria and two modified approaches that use spleen stiffness measurements to identify patients who may be candidates for nonselective beta-blocker therapy without invasive testing.
Current Baveno VII criteria use liver stiffness and platelet count to assess portal hypertension but leave 40% to 50% of patients with inconclusive results, according to the investigators. To improve accuracy, they developed a machine learning model that combines liver stiffness, spleen stiffness, platelet count, Child-Pugh score, age, sex, and cause of liver disease. The model can be used with vibration-controlled transient elastography, two-dimensional shear wave elastography, and point shear wave elastography.
The derivation cohort included 1,093 patients from 17 centers, including 943 used to develop the model and 150 used for internal validation. An independent external validation cohort included 342 patients from seven additional centers. More than half of the patients in every cohort had CSPH confirmed by invasive HVPG measurement. Compared with the derivation cohort, the external cohort included more patients with metabolic dysfunction-associated steatotic liver disease, allowing the investigators to test the model across a broader range of liver diseases.
Among the five machine learning models tested, the random forest model performed best. In the external validation cohort, it achieved an area under the receiver operating characteristic curve of 0.91. Using predefined cutoffs, an ELM score of 0.45 or lower ruled out CSPH with a 90% negative predictive value, while a score of 0.60 or higher ruled it in with a 96% positive predictive value.
The ELM score reduced the proportion of patients with inconclusive results to 12%, compared with 48% using standard Baveno VII criteria. That was also lower than the 39% seen with the Baveno VII dual spleen stiffness model and the 20% seen with the single spleen stiffness model. Overall, the ELM score correctly classified 88% of patients with CSPH as having the condition, compared with 58% using standard Baveno VII criteria.
The ELM score performed consistently across all three elastography techniques. Depending on the imaging method, 7% to 12% of patients had inconclusive results, compared with about 45% to 56% using conventional criteria. The model also performed well across different causes of liver disease. Gray-zone rates ranged from 5% in autoimmune liver disease to 11% in alcohol-associated liver disease and 14% in patients with mixed or other causes. Obesity increased the rate of inconclusive results but had little effect on overall diagnostic accuracy, the investigators reported.
The investigators also compared the ELM score with the ANTICIPATE and NICER scores in patients who underwent vibration-controlled transient elastography. The ELM score showed similar accuracy in identifying patients with CSPH while reducing the proportion of inconclusive results to 12%, compared with about 41% for both existing scores.
In an exploratory analysis, the investigators assessed whether the ELM score could predict future hepatic decompensation in 190 patients who were followed over time. No patients with an ELM score below 0.60 developed hepatic decompensation at three or five years of follow-up, suggesting the model may help identify patients at very low risk. However, the investigators emphasized that this finding needs to be confirmed in prospective studies.
For physicians, the findings suggest the ELM score could improve selection of patients for nonselective beta-blocker therapy while reducing the need for invasive HVPG measurement. Patients with scores of 0.60 or higher could be considered for treatment after physician review, while those with scores of 0.45 or lower might avoid invasive testing. Patients with intermediate scores would still require confirmatory testing, such as HVPG measurement or upper endoscopy.
Dr. Colecchia and colleagues noted several limitations. Most cohorts were retrospective, raising the possibility of bias, and differences among elastography platforms may have affected performance. The model also requires prospective validation in larger, more diverse populations and regulatory approval before routine clinical use.
The study received no specific funding. Several authors reported relationships with pharmaceutical and medical device companies.
Expert Insight
Dr. Colecchia discussed the study’s potential clinical implications with GI & Hepatology News.
How do you envision the ELM score changing the current approach to identifying patients who should receive nonselective beta-blockers?
Dr. Colecchia: The ELM score could improve non-invasive risk stratification for clinically significant portal hypertension by integrating liver and spleen elastography measurements with biochemical data and key clinical variables, including age, sex, and disease etiology. Importantly, it is the first pan-elastography risk-stratification score designed to harmonize measurements across different elastography modalities. By combining these parameters, ELM substantially reduces the diagnostic gray zone and provides clearer, actionable thresholds. This could help clinicians identify patients in whom nonselective beta-blocker (NSBB) therapy can be initiated with greater confidence, while allowing others to avoid or defer invasive testing.
Which patients are most likely to benefit from the ELM score compared with the current Baveno VII criteria?
Dr. Colecchia: Patients who fall within the diagnostic “gray zone” under the current Baveno VII criteria are likely to benefit most. The ELM score can reclassify many of these indeterminate cases into clearer rule-in or rule-out categories. This may enable more timely treatment decisions while reducing unnecessary procedures, such as hepatic venous pressure gradient measurements or endoscopic surveillance.
What additional validation is needed before the ELM score can be adopted in routine clinical practice?
Dr. Colecchia: ELM is currently intended for research use only. Although ELM has undergone multicenter external validation, further independent, real-world replication across additional centers, patient populations, and elastography devices will be essential. These studies should confirm that the model’s diagnostic performance and proposed cutoffs remain consistent across diverse clinical settings. This should be followed by prospective implementation studies evaluating outcomes such as reduction of the diagnostic gray zone, time to NSBB initiation, use of hepatic venous pressure gradient measurement, and safety signals—including hypotension and renal adverse events. Health-economic outcomes should also be assessed, particularly the number of invasive procedures avoided per correctly classified patient.