Digital Pathology Podcast

204: Assessing interstitial fibrosis and tubular atrophy in kidney biopsies artificial intelligence versus humans

Aleksandra Zuraw, DVM, PhD Episode 204

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 18:47

Send us Fan Mail

Paper Discussed in this Episode:

Assessing interstitial fibrosis and tubular atrophy in kidney biopsies artificial intelligence versus humans. Farris AB, Zukić D, Solez K. Current Opinion in Nephrology and Hypertension. March 16, 2026.

Episode Summary: In this journal club deep dive on the Digital Pathology Podcast, we explore the intense debate over quantifying chronic kidney disease progression. We unpack a fresh 2026 study comparing artificial intelligence to human pathologists in assessing interstitial fibrosis and tubular atrophy. If top experts can't agree on a diagnosis due to human subjectivity, can an AI trained on their imperfect data provide a better standard? We explore what happens when pixel-perfect machines clash with nuanced human medical judgment.

In This Episode, We Cover:

The Clinical Stakes of Kidney Scarring: Why interstitial fibrosis (the scarring of tissue spaces between filtering units) and tubular atrophy (shrinking and collapsing functional tubes) are the primary surrogate measures for tracking chronic kidney disease. We discuss how a mere 10% diagnostic variance can drastically alter a patient's medication regimen, dialysis prep, or transplant eligibility.

The Flaw in the "Gold Standard": We break down the "interobserver variability" problem—why two highly trained, board-certified pathologists can look at the exact same biopsy slide and give two completely different mathematical assessments of the damage.

How the AI Actually Works (Mapping the Neighborhood): A look at "indirect assessment through kidney compartment segmentation," where the AI acts as a digital surveyor. It identifies cellular fences like glomeruli and tubules, establishing microscopic "zoning laws" before it begins counting the damaged tissue.

The Proofreader vs. The Literary Critic: Why studies show a persistent "lack of complete concordance" between human and machine. We discuss how AI hyper-focuses on mathematical pixel intensity and mistakes physical slide artifacts (like a folded piece of tissue) for severe disease. Meanwhile, human pathologists act as "literary critics," easily filtering out the visual noise using clinical context.

The Humans + AI Synergy: The ultimate endgame isn't replacing pathologists, but combining the tireless mathematical consistency of AI with the complex contextual reasoning of humans to create a highly advanced co-pilot system.

Key Takeaway: The lack of perfect agreement between AI and human pathologists isn't a failure, but rather evidence that they perform fundamentally different types of analysis. AI excels at tedious, reproducible quantification that eliminates human visual fatigue, but it lacks contextual judgment. By adopting a "humans + AI" workflow, the medical field can stabilize crucial kidney measurements and elevate the pathologist to a true diagnostic synthesizer, ultimately leading to more effective patient care

Support the show

Get the "Digital Pathology 101" FREE E-book and join us!

Imagine basing like a lifealtering medical treatment on a precise measurement, right? Yeah.


But then you realize that the world's absolute top experts cannot actually agree on what that measurement is.


Yeah. It's it's honestly an uncomfortable truth about how we quantify disease,


right? Because you have the tissue, you have the microscope, you have decades of clinical experience, and yet the final number comes down to well, a subjective interpretation.


Exactly. I mean, we expect pathology to be binary like a bone fracture on an X-ray. But when you are looking at complex cascading organ failure, the landscape is just incredibly murky.


It's a terrifying thought for a patient.


Oh, absolutely. You are trying to put a hard mathematical number on a biological process that is frankly chaotic and uneven.


And that is exactly the mystery we are diving into today. Welcome trailblazers to today's deep dive on the digital pathology podcast. We have a highly focused journal club style analysis lined up for you.


Yeah, we're looking straight into the literature that is ly reshaping the bedrock of your field right now


and it is incredibly fresh. Published just two days ago on March 16th, 2026 in the journal current opinion in nefrology and hypertension. The authors are Alton B. Ferris, June Zukich, and Kim Solles.


The timing on this publication is just crucial. I mean, departments everywhere are currently debating the exact role of automated image analysis and this paper hits right at the center of that tension.


The paper is titled Assessing Interstatial Fibrosis and Tubular atrophy in kidney biopsies, artificial intelligence versus humans.


Quite a mouthful, but incredibly important.


Definitely. So, our mission today is to explore the intense ongoing study of how artificial intelligence compares to and potentially augments human pathologists in these highly critical kidney assessments.


Right? We are going to look at why humans struggle with this, how the machines are attempting to solve it, and well, what happens when the two completely disagree.


Okay, let's unpack this. Why did Ferris and his team write this paper in the first place because looking at the author's summary, they highlight a very specific clinical problem driving this massive push for AI.


The catalyst for all of this research is human limitation really, especially in the face of immense visual complexity.


Yeah.


To understand the push for computerized methods, you have to look closely at the specific diagnostic features the paper focuses on, which are interstatial fibrosis and tubular atrophy.


Right? And in nephrology, these are not just minor footnotes on a pathology. report. They are widely used as the primary surrogate measures for chronic kidney disease progression.


Exactly. They are the core indicators.


Let's ground this for a second because we use the terms fibrosis and atrophy constantly, but it helps to visualize what is actually happening to the tissue.


Yeah,


interstatial fibrosis is essentially the scarring of the tissue spaces between the kidneys filtering units.


Spot on. And tubular atrophy happens right alongside it. It's when the kidney's filtering tubes literally shrink, collapse, and just lose their functional capacity


like pipes getting crushed.


Yeah. That structural breakdown is the absolute hallmark of a failing kidney. So when a pathologist looks at a biopsy under the microscope, usually treated with specific chemical stains to make the scarred tissue highlight in blue or green, they are trying to estimate the total percentage of the biopsy that has been taken over by this scarring and shrinking.


Right? But the problem the paper points out like right in its opening summary is that interobserver variability among human pathologists has been consistently demonstrated when assessing these exact features.


Meaning, you know, you can take the exact same kidney biopsy slide, put it under a microscope, and two highly trained, board-certified pathologists might look the exact same stained tissue and give you two distinctly different assessments.


Two different numbers for the exact same slide.


Yeah, that variability is heavily documented. Human eyes and brains are incredible at recognizing patterns, but they are notoriously poor at consistently quantifying area or volume across a complex irregular surface.


It makes me think of taking a classic car to two expert mechanics to get an estimate on the rust damage on the undercarriage.


Oh, that's a great analogy,


right? Both mechanics look under the car and instantly agree on the pathology. They say, "Yes, this is bad. The structural integrity is failing."


The pattern recognition is there.


Exactly. But if you demand an exact mathematical percentage of rust coverage, one might estimate 30% and the other might say 40%. They are human beings trying to visually estimate a chaotic spread.


Yeah, but in a clinical setting, a 10% variance isn't just an academic debate, you know.


Far from it. That 10% variance completely changes the clinical stakes for the patient you are treating.


Oh, absolutely. In chronic kidney disease, the difference between a 30% and 40% fibrosis reading might determine whether a nefologist aggressively adjusts a medication regimen


or begins preparing the patient for the reality of dialysis. right? Or alters their eligibility status for a kidney transplant entirely. The stakes of that measurement are absolute, which makes the subjective variability a profound clinical problem.


But wait, this brings up a massive question for me, and I want to push back on the underlying premise of this technology.


Sure, go ahead.


If human assessment is the gold standard we currently use to track chronic kidney disease, but the paper explicitly admits there is significant interobserver variability among those humans, are we essentially building our new AI models on a shaky foundation.


That is a very valid concern.


I mean, if the AI requires thousands of human-graded slides to learn from, but the pathologists providing the grades do not completely agree with each other, how does the AI ever know what the definitive ground truth actually is?


What's fascinating here is how the authors tackle that because it is perhaps the single biggest unresolved debate in digital pathology right now. They do not shy away from it.


It's a real paradox.


It is a fundamental structural problem in applying machine learning to subjective medical fields. The algorithms are indeed being trained on noisy, imperfect human data.


So, how is that useful?


Well, the goal of bringing an AI, as framed by the study, isn't necessarily to magically discover some divine perfect ground truth that humans missed. The immediate goal is to introduce a level of standardized reproducible measurement.


Ah, I see.


Even if the AI's baseline is a synthesized average of human opinion, The AI will apply that standard with mathematical consistency across thousands of slides.


Something a human eye experiencing visual fatigue simply cannot sustain.


Exactly. So the consistency itself is a huge upgrade.


And because human variability is the core issue, the natural next step in the paper is examining how the computers are actually attempting to execute that consistency.


Right. The authors outline the recent findings on how computerized assessment operates, distinguishing between two main AI approaches,


which are direct inter stitial fibrosis measurement and indirect assessment through kidney compartment segmentation.


The distinction between those two approaches is really vital. Drought measurement is essentially training the computer to look at the entire image and quantify the specific colors and textures associated with scarring.


But the indirect assessment through kidney compartment segmentation that seems like a much more sophisticated way of mimicking how a human actually evaluates tissue architecture.


Oh, it is. But the phrase indirect assessment through compartment segmentation is well it's pretty dense jargon.


Let's break down the actual mechanism for the trailblazers listening.


Yeah.


How does the AI achieve this? Because it isn't just looking at a JPEG and counting blue pixels.


No, it is mapping the structural geography of the kidney first. In this segmentation approach, the AI looks for the microscopic boundaries,


the cellular fences, if you will.


Exactly. It methodically identifies the dense cellular clusters of the glumberi which are the main filtering hubs. Then it establishes the borders of the tubules.


Okay. So once it has successfully mapped out the normal architecture and segmented those compartments, what does it do next?


Well, it looks at the spaces in between the interstitium to calculate the exact volume of fibrous tissue or to measure exactly how much a specific tubule has atrophied compared to its healthy baseline.


So it is establishing the neighborhood zoning laws before it starts counting the damaged houses.


That makes a lot of sense. Yeah, that's exactly what it's doing.


But the text also uses a phrase that caught my eye. While discussing these computerized assessments, it evaluates AI alongside what it calls handcrafted methods.


Right?


In a field dominated by neural networks and deep learning, why are the authors bringing up handcrafted computer algorithms?


It provides necessary historical and technical context. Before deep learning AI models became the dominant force, you know, the ones that teach themselves to recognize patterns based on massive data sets, computer scientists use handcrafted methods. traditional software.


Yeah, these were traditional rule-based algorithms where a human programmer manually wrote the explicit mathematical rules for the software. They would code the exact pixel intensity, the exact color threshold, and the specific edge detection parameters the computer should use to identify fibrosis.


So, the human writes the mathematical recipe and the computer just bakes the cake faster.


Exactly. But with modern AI, the computer is given a thousand cakes and figures out the recipe on its own.


That is a very way to put it. And the paper evaluates how both of these computerized approaches, the rigid hand-crafted algorithms and the modern pattern recognizing AI stack up when placed head-to-head against human pathologists,


which brings us to the ultimate clash. Really,


we have established these high-tech methods. The AI is methodically segmenting the kidney compartments, finding the cellular fences, and counting with pixel level precision.


And the computers do not need coffee. They do not get tired after reviewing their 50th slide of the morning and they do not have subjective bias.


So when the authors looked at the studies comparing AI to humans, did the machines perfectly solve the variability issue?


They did not.


Wait, really?


Yeah. When the authors synthesized the current research comparing computerized methods and human experts, they report a distinct, highly critical finding, a persistent lack of complete concordance.


Here's where it gets really interesting. A lack of complete concordance. Let me challenge the interpretation of that. result for a second.


Go for it.


If AI is mathematically precise, if it is methodically segmenting these compartments based on vast amounts of training data, why isn't there complete concordance? If the human is the one suffering from documented interobserver variability, maybe the human gold standard is simply disagreeing with a potentially more accurate computer.


That is exactly what everyone wonders at first,


right? Why do we assume the lack of concordance is a failure on the AI's part?


It is incredible. tempting to assume the machine is the ultimate orbiter of truth because its calculations are absolute. But the author's caution against that exact assumption.


Oh, okay.


The text does not declare AI the undisputed winner just because it is precise. Precision is not the same as accuracy, especially in biology.


That's a huge distinction.


It really is. Despite the immense quantitative capabilities of these computerized methods, the authors explicitly note that studies still show the persistent value of human assessment in many circumstances.


Persist value.


Yeah.


Meaning the nuanced human eye is catching biological realities that the pixel perfect machine is completely dropping.


Let's look at the mechanism of why that happens. Consider the physical reality of a biopsy. You are taking a three-dimensional piece of tissue, slicing it incredibly thin with a microone blade,


mounting it on a glass slide, bathing it in chemicals,


right? And during that physical process, artifacts happen. A slice of tissue might fold over on itself slightly,


like a Sprinkle in a piece of tape


exactly like that. To a computerized method, whether handcrafted or deep learning, that folded tissue presents as a dense, over overlapping, highly concentrated cluster of cells in stain.


Oh, I see where this is going.


Yeah. The algorithm following its mathematical training might rigidly interpret that dense area as severe interstatial fibrosis or profound tubular collapse. It registers as a critical disease state.


But a human pathologist looks at that exact same cluster. M


and instantly recognizes the visual signature of a tissue fold.


Exactly. They understand the context of the preparation. They mentally subtract that artifact from their assessment without a second thought.


Wow. So the AI flags it as endstage disease, but the human knows it is just a folded slide.


It's the difference between a hyperfocused proofreader and a literary critic.


That is a brilliant way to phrase it. The AI is the proofreader. It can scan a 500page book and tell you the exact mathematical number of vowels. in a fraction of a second


and it will never miscount a vowel.


But the human pathologist is the literary critic. They actually understand the plot, the context, and the nuance of the story being told by the cells.


Yeah, the proof reader might flag a metaphor as a factual error while the critic understands the deper meaning. What's fascinating here is how perfectly that analogy captures the author's findings regarding concordance.


So the lack of complete concordance is not a sign that AI is useless, nor is it a sign that humans are obsolete.


No, it is evidence that they're performing two fundamentally different types of analysis. The computer excels at the tedious reproducible quantification that humans are notoriously inconsistent at.


But the machine completely lacks the nuanced medical judgment required to synthesize complex overlapping morphological exceptions. So we have a fascinating tension here. Humans have interobserver variability making them an imperfect subjective gold standard for tracking chronic kidney disease progression. Right.


But we also have highly advanced AI and handcrafted computerized methods that lack complete concordance with the humans because they trip over tissue artifacts and lack clinical context.


That's the catch 22.


So what does this all mean? Where does this actually leave the trailblazers listening today? If I am a pathologist sitting in a lab right now hearing about AI segmenting kidney compartments, I might be worried that I am training my own automated replacement.


Well, what is the actual end? game. According to Ferris, Zukich and Solles,


that's what I want to know.


If we connect this to the bigger picture, the authors are very clear in their summary about the trajectory of the field. They note that computerized methods are unequivocally showing increased application for a wide variety of clinical and hystopathologic parameter assessments.


Okay, so the technology is expanding its footprint in the lab.


It is, but they immediately follow that with a crucial caveat. Additional work is needed to fully integrate these computerized methods into routine pathology practice. The endgame, as they explicitly state, is not replacement. It is synergy.


It is the integration of the proofreader and the critic.


Exactly.


The authors are advocating for a workflow where the machine processes the vast, overwhelming amounts of visual data. It instantly maps the cellular fences, segments the compartments, and provides a perfectly reproducible, mathematically consistent baseline measurement of the interstatial fibrosis.


It removes the visual fatigue entirely. But the human remains the definitive diagnostic authority.


Like a highly advanced co-pilot system, the machine processes the vast data, but the human commands the flight.


That is the exact collaborative tool set the paper envisions. The human pathologist reviews the AI segmented data, provides the clinical context of the patient's history, catches the tissue folds and morphological artifacts that the algorithm misinterpreted, and then makes the final diagnostic call.


And the authors state that this specific combination, humans plus AI, may ultim provide enhanced analysis for more effective patient care


and more effective patient care is the entire objective.


Absolutely. If we can merge the mathematical consistency of the AI with the nuanced judgment of the human, we can finally stabilize that surrogate measure for chronic kidney disease.


We stop arguing over whether a biopsy is 30% or 40% damaged and we start making faster, more confident decisions about medication, dialysis, and transplant timelines.


It elevates the pathologist from a manual pixel counter to a true diagnostic synthesizer.


The human is freed up to do what the human brain actually excelled at, complex contextual reasoning.


Let's take a quick look back at the journey we have been on today. We started with the foundational problem driving this research. Human pathologists, despite their brilliance, suffer from inner observer variability when visually assessing interstatial fibrosis and tubular atrophy,


which is a massive clinical hurdle since those features are our primary surrogate measures for tracking the progression of chronic kidney disease.


We then explored how computerized methods are stepping in to address this variability, moving from traditional handcrafted algorithms to complex AI models that map out the kidneys architecture through indirect compartment segmentation.


But we also discovered that the machines are not a flawless out-of-the-box solution. The studies show a distinct lack of complete concordance between AI and humans.


Right? Proving that algorithms still struggle with tissue artifacts and clinical context. The nuanced human eye still holds immense persistent value.


Which ultimately leads us to the vision outlined by the authors. A future where the standard of care is defined by humans plus AI working in tandem to deliver a level of diagnostic accuracy that neither could achieve independently.


We want to thank all of you trailblazers for joining us on this deep dive here on the digital pathology podcast. We know your time in the lab and the clinic is incredibly valuable


and staying on top of the rapidly shifting literature is no small feat. But as we wrap up today's discussion on the messy, fascinating intersection of tissue analysis and artificial intelligence, we want to leave you with something to chew on.


This raises an important question based on everything we've unpacked today regarding the discrepancy between human and machine vision.


Right? If AI is capable of evaluating pixel level data across millions of biopsies and it consistently lacks concordance with humans on certain complex features, could the AI eventually begin to identify entirely new microscopic pattern? patterns in interstial fibrosis that human science hasn't even defined yet.


And if so, if the machines start seeing biological truths that our eyes literally cannot perceive, how long until that humans plus AI dynamic forces us to completely rewrite the very definition of chronic kidney disease progression?


Until next time, keep blazing those trails.