Liver stiffness readings varied nearly 20% between visits in MASLD study

Share

Changes in noninvasive test results for metabolic dysfunction-associated steatotic liver disease (MASLD) may need to exceed expected measurement variability before they are considered clinically meaningful, according to a study published in the Journal of Hepatology.  

The findings provide reproducibility benchmarks for blood-based, elastography-based, and composite biomarkers used to monitor MASLD.

The investigators analyzed biomarker data from the pre-randomization screening period of the phase 2 MIRNA trial, which evaluated ervogastat and clesacostat in patients with biopsy-confirmed at-risk metabolic dysfunction-associated steatohepatitis, or MASH. The reproducibility analysis included two cohorts. The screening analysis set included 803 participants across the full MASLD spectrum. The randomized analysis set included 255 participants with histologically confirmed at-risk MASH, defined as a MASLD Activity Score of 4 or higher with F2-F3 fibrosis.

The study was conducted at 198 centers in 11 countries, with biomarkers measured repeatedly before participants received pharmacologic intervention. The analysis included routine laboratory markers, simple fibrosis scores such as FIB-4 and the aspartate aminotransferase-to-platelet ratio index, measures derived from vibration-controlled transient elastography (VCTE; FibroScan), composite panels, and circulating markers of necroinflammation or extracellular matrix turnover.

Reproducibility was assessed using the within-subject coefficient of variation (wCV) and the reproducibility coefficient (RDC). The wCV expresses within-subject variability as a percentage relative to the mean of a participant’s repeated results. The RDC is the value below which the relative change between repeat measurements would be expected to fall in 95% of cases, assuming the participant’s disease state had not changed. Together, these measures helped estimate whether short-term changes were more likely to reflect measurement variability or true disease progression or treatment response.

In the broader screening cohort, liver stiffness measurement by VCTE showed substantial short-term variability, with a wCV of 19.9% and an RDC of 55.2%. Variability was higher in the randomized at-risk MASH cohort, at 23.0% and 63.8%, respectively.

Controlled attenuation parameter showed lower variability among VCTE-derived measures, with wCVs of 8.8% in the screening cohort and 8.7% in the randomized cohort. Among all biomarkers assessed in the randomized cohort, the enhanced liver fibrosis (ELF) score showed the lowest variability, with a wCV of 3.4% and an RDC of 9.3%.

Markers related to steatohepatitic activity were more variable. In the randomized cohort, cytokeratin-18 (CK18) fragments M30 and M65 had wCVs of 30.7% and 27.8%, respectively. Their RDCs were 85.1% and 77.0%. The authors wrote that higher variability among steatohepatitis-related biomarkers “may reflect the more dynamic nature of steatohepatitis.”

Type I and Type II errors with the presence of wCV/RDC. (A) and (B) are conceptual representations that illustrate how the differences between adopting thresholds defined using wCV and RDC respectively could affect the performance of a biomarker to detect a biologically meaningful change. In general, adopting RDC thresholds over wCV will minimise FP (i.e. increased specificity) but at the cost of an increase in FN (i.e. reduced sensitivity). (C) and (D) show the relative range of “no biological change” biomarker variation across a range of biomarkers in the RAS cohort. The upper/lower bounds of biomarker variation thresholds using within-subject %CV (C) and RDC (D) are shown. TN represent individuals without clinically meaningful changes, while TP are those with clinically meaningful changes. FN indicate missed true changes, and FP reflects incorrectly identified changes. These results cannot be directly compared across biomarkers without accounting for differences in biomarker magnitude. Reproduced from Au TY, et al. J Hepatol. Published online Aug. 3, 2026. doi:10.1016/j.jhep.2026.07.029. Licensed under CC BY 4.0.

For clinicians, the study suggests that small changes in noninvasive test results should be interpreted cautiously during longitudinal MASLD monitoring, particularly when they fall within expected test variability. A modest shift in liver stiffness, liver enzymes, or composite scores may reflect biological or technical variability rather than true disease progression or treatment response.

The authors wrote that the lack of robust reproducibility data has hampered longitudinal use of noninvasive tests by creating “uncertainty about what magnitude of biomarker change indicates genuine disease progression.” They added that the same evidence gap has limited regulatory adoption of noninvasive tests as surrogate endpoints in drug development.

The results may also inform therapeutic trials in MASH, where investigators need to distinguish true treatment response from background biomarker variability. The authors cautioned that reproducibility thresholds should be used to interpret change within a given biomarker, not to compare one test with another. They also suggested that paired changes in complementary measures, such as elastography and a circulating fibrosis marker, could offer a balanced approach, although this strategy still needs to be evaluated in clinical trials.

The authors noted several limitations. The analysis assumed that participants’ underlying disease did not change meaningfully during the screening period. They said this was plausible given the short observation window and generally slow progression of MASLD.

The authors also noted that MIRNA was not designed primarily to assess biomarker reproducibility, and that the analysis was conducted under clinical trial conditions, with centralized laboratories, standardized procedures, and sponsor quality control. They wrote that reproducibility may differ in routine practice.

The authors concluded that the findings provide actionable reproducibility data for widely used noninvasive tests.

The MIRNA trial was funded by Pfizer. Several authors reported employment, stock ownership, consulting fees, lecture fees, research funding, travel support, royalties, advisory board participation, or other relationships with pharmaceutical, diagnostics, and medical technology companies.