
Reliability 4.0 · 2026-06-16 · 17 min
Key moments - from our scoring
Substance score
47 / 100
Five dimensions, 20 points each
Fresenius Medical, a kidney disease treatment company producing hemodialysis and peritoneal dialysis consumables, has adopted a structured Rapid Problem Solving approach for equipment breakdowns since October 2024. Juan Mendoza explains how RPS differs fundamentally from traditional RCA: it's a time-boxed methodology (24-hour completion, 5-day investigation window) designed to resolve acute equipment failures quickly on the production floor using cross-functional teams, while RCA and eight-discipline analysis are reserved for chronic, complex process defects requiring deeper investigation. Mendoza advocates treating roughly 80% of issues with RPS basics - focusing on containment, GEMBA (shop floor) investigation, and rapid implementation - while reserving the remaining 20% for intensive root cause work. He emphasizes top-down implementation, connecting improvements to OEE metrics (specifically unplanned downtime reduction), and highlights emerging use of AI-powered vibration monitoring via Traction for condition-based maintenance. The episode's core insight: before investing in sensors and AI, master the fundamentals of alignment, lubrication, fasteners, and balance (the FLAT concept) - an approach applicable to any manufacturing operation seeking reliability gains without over-engineering solutions.
RPS is a rapid, time-boxed investigation (24-hour initiation, 5-day completion) focused on acute equipment breakdowns, while RCA and eight-discipline analysis are deeper methodologies for chronic, complex process issues; RPS trades some thoroughness for speed to prevent issues from recurring quickly on the shop floor.
1) Containment action to restore equipment to production first, 2) Begin RPS within 24 hours, 3) Conduct investigation on the gemba (shop floor) with a cross-functional team, 4) Complete investigation within five days, 5) Implement actions and follow-up to prevent recurrence.
Implementation must cascade from top to bottom - managers must first believe in and practice RPS before cascading to technicians, which prevents technicians from dismissing it as a quality issue imposed on them.
The primary metric is OEE (Overall Equipment Effectiveness), with maintenance-specific focus on unplanned downtime reduction; other organizations also track Mean Time Between Failures (MTBF) as a secondary metric.
FLAT stands for fasteners, lubrication, alignment, and balance; these four fundamentals address approximately 80% of equipment issues and should be mastered before implementing advanced monitoring technologies.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode delivers a handful of genuinely useful operational details - specific thresholds, a dual-track framework, the FLAT acronym - but meaningful content is front-loaded into roughly 12 minutes after stripping out the boxing small-talk and outro promo. Insight-per-minute is adequate, not dense.
any hours or interruption above 45 minutes, we have to do an RPS. And we have been implementing this since October 2024
I would think that 80% of the issues will be solved through the basics, such as the RPS. The other 20, it requires more deep investigation
The guest himself acknowledges the content is established lean/TPM practice ('it's not kind of the brain-bent in the world. These are really all tools that all the companies use'), and the frameworks - RPS, 8D, PDCA, OEE - are textbook. The FLAT acronym from an unnamed consultant is the only mildly fresh framing.
it's not kind of the brain-bent in the world. These are really all tools that all the companies use
the RPS is more focused on equipment breakdowns...The RCA or the depth that we use here is more related with the process
Juan Mendoza is a genuine practitioner - Director of Manufacturing Engineering at a regulated medical-device manufacturer with 25 years of hands-on experience across Alcoa and Fresenius - not a thought-leader or conference circuit speaker. His depth is real but he is a mid-tier operator rather than a transformational figure, which caps the ceiling.
I learned this, as I mentioned, about 25 years ago, in a company in Mexico named Alcoa, and they're based on Toyota production system
my role here is director for manufacturing engineering. So I'm in charge for everything related with maintenance and manufacturing engineering
There are usable concrete details - specific time thresholds (45 min, 24 hrs, 5 days), the $100k aluminum chamber case study with a root cause traced to operator-induced overheat, named vendors (Alcoa, Traction, Jiri), and October 2024 as implementation date - but most examples stop short of quantified outcomes or before/after data.
that equipment price was above $100,000
we have to complete this RPS in no longer than 24 hours...we have to finish the investigation in no longer than five days
The host asks a couple of genuinely useful follow-ups (enforcing culture change, RPS vs. RCA distinction, trading accuracy for speed) but opens with several minutes of biographical small-talk and boxing anecdotes, lobs generic innovation and 'final word' questions, and never challenges any claim or probes for outcome data.
So are you effectively saying you're trading a little bit of maybe accuracy for speed?
How did you enforce those standards?
Computed from the transcript - who did the talking, and the words that came up most.
Sometimes the perfect RCA becomes the enemy of solving the problem. In this episode of Reliability 4.0, Juan Mendoza from Fresenius Medical Care shares how his team uses Rapid Problem Solving (RPS) to respond to equipment failures faster, reduce repeat downtime, and build stronger problem-solving habits across manufacturing teams. Juan explains why most equipment issues don’t need weeks of investigation - and how setting clear rules around speed, ownership, and execution helps teams stay focused on solving problems instead of endlessly analyzing them. We discuss: Why every unplanned downtime over 45 minutes triggers an RPS. How “containment first” keeps production moving before investigation starts. Why rapid investigations must be completed within five days. The difference between rapid problem solving and deep RCA investigations. Why most reliability improvements still come from mastering the basics. If your organization struggles with recurring downtime or slow investigations, this episode offers a practical framework for building a faster reliability culture. This episode is available on: YouTube: Apple Podcasts: Or wherever you get your podcasts!
Transcribed and scored by The B2B Podcast Index.
Hi, everyone. This is Sebastian Traeger, Managing Director, and welcome back to another episode of Reliability 4.0. As always, we're on a mission to make the world a more reliable place.
And today, we are joined by Juan Mendoza with Fresenius Medical. We are excited to have him. Juan, thank you so much for joining us to share your work and your insights with us. Well, thank you, Sebastian, for the invitation.
And I'm very happy to be here in your podcast. podcast. So you were telling me you're in Knoxville. Tell us how you ended up coming to work at Fresenius and how you came to be in Knoxville.
Yes, I'm in this company for three years before this company was on St. Louis, Missouri. And this was an opportunity for grow up professionally and as well the area. So I really like Knoxville.
There is a lot of places for outdoor activities and my family is happy as well here. That's great. Well, I'm glad to hear that. Now, is that like a boxing belt above your shoulder?
Yes, there is a boxing belt, but it's not for boxing. We just received this recognition about two weeks ago for a company named Jiri. They create a software for automated electronic work instructions. and we received it as the key customer for 2025.
Excellent. All right. Well, you look like you could have been a boxer. So I just had to ask the question to see if that was in your background at all.
My son actually is practicing boxing for almost three years. Have you ever boxed against him? No. At the beginning, I might try, but at this point, I will not try.
I've only boxed once in my life And it was, you know, kind of on a whim It was with some friends And we had boxing gloves We're like, oh, let's put them on and box It is the most exhausting sport Oh my goodness I was like, I could not believe how exhausting that was To do that That's true So that's good That's pretty neat that your son's doing this Well, tell us a little bit about Fresenius And tell us about your role there Yes, Fresenius is a company that provide services and products for kidney disease, pretty much for the two main treatments, which is hemodialysis and peritoneal dialysis.
In our case, we produce the consumables, which is a solution for peritoneal dialysis and a vicar for hemodialysis. And my role here is director for manufacturing engineering. So I'm in charge for everything related with maintenance and manufacturing engineering. That's obviously a very big job, being in charge of maintenance and manufacturing.
So you're responsible to make sure all the product gets out and responsible for the team that does the operating and as well as the maintenance. That's correct. That's great. Well, that's good to hear.
And you had recently done a talk on using rapid problem solving tools to diagnose equipment failures. What are some of the failures in your business that you see? and what would you say are the best practices for bringing rapid problem solving to bear? Yes, pretty much everything related with equipment breakdowns or unplanned downtimes.
So I learned this two times ago, about 25 years ago, in a company in Mexico named Alcoa, and they're based on Toyota production system. They call it ATRI, and they decide to use this tool for every breakdown or interruption production over 15 minutes. That was very tight. In our case here in Resenius, we established this as a 45-minute threshold.
So any hours or interruption above 45 minutes, we have to do an RPS. And we have been implementing this since October 2024 and with really good results. So the team is engaged, And I highly recommend this tool for all the organization. Can you give us an example of what, you know, just kind of take us from problem discovery to kind of the process you put that through in order to solve a problem?
Yes sure Well actually there are five rules that we established for this methodology The first thing is the containment action So we have to make sure that the equipment is back again in production That's the priority. So we cannot start doing an RPS while the equipment is still down. So we should focus first on the containment. So after the equipment is back in operation, the second rule is that we have to complete this RPS in no longer than 24 hours.
Let's say that we had a leak in our pipe. We focus on fix the leak, but then before 24 hours, we have to start the RPS. And the other rule is that we have to do this in the GEMBA. So we have to go to the floor where this happens.
I think a team, that's the other rule. We have to have this in a team. And then we have to find the root cause for this issue. The reason to have this in the floor, because there is all the evidence is there and having a group of people who have different perspective.
The other rule is that we have to finish the investigation in no longer than five days, because if we don't complete this in five days, that's not rapid. That's why it is called a rapid problem solving. And then we have to implement all the actions and follow up. And the benefits of this is avoid the box coming up again.
So if we don't, if we just focus on the repair, there is high possibility that this will happen again. Do you, and it sounds incredible, that process and kind of the standards that you guys set. How did you enforce those standards? What I assume people don't like to say, hey, I got to do these five steps.
So how did you kind of bring about that culture change? Yes, that's a good question, Stens, for asking. One important thing that is kind of a priority for the implementation is this needs to go from top to bottom. So we have to start first with the managers in order for them to believe and see the benefits of the tool.
So that's how we started in October. I was the first who started the process. So when this happened, I jumped as a facilitator for the RPSs and then cascade with the managers. So then the manager started doing the process through the time, and now they are cascading with the technician.
It's not something kind of that quick, but needs to go top down. Otherwise, if you right away assign this to a technician, probably they will connect this with such quality. And they said, no, this is not for us. This is more a quality issues.
But that's the best way to cascade from top to bottom. So is that a process you brought with you from a different organization? Or did you have someone, again, top down say to you, this is what I want you to do? Yes, I learned this, as I mentioned, about 25 years ago, and I have been using for all the organizations that I have been working here.
Hidden in Fresenius, and it's not kind of the brain-bent in the world. These are really all tools that all the companies use, but the methodology in regards of the implementation and the rules to execute it is kind of a key for having a successful outcome. Okay, that's good. What metrics do you focus on then?
So, you know, when you talk about bringing this to bear and seeing results, what are those metrics that you're seeing to indicate that problems are getting solved, they're not recurring and kind of making progress on what you're hoping to? Yes, sure. Well, here in Fresenius, our main metric is OEE. We connect all the improvements to the OEE.
That's where we want to see the results. And more specifically, we break the OEE for maintenance is the unplanned downtime. So we have a threshold for unplanned downtime, and that's where we are seeing the improvement. And yes, of course, we saw the improvements through the time.
In other organizations, here we are not doing this yet, but we will do in the near future. Mean time between failures is another metric. For example, in Alcoa about 25 years ago, we used that metric. Here we are not using that yet, but that the plan for the near future Excellent So our company does a lot around root cause analysis and you talking about you know not RCA but RPS How are these different or similar in your mind?
Well, the difference, RCA is part of the RPS. The difference, I will say, this is a quick investigation to apply solutions right away. It is not, I would say, 100% effective or efficient, but it gives you the opportunity to complete the actions quick. There are other problems that are more complex, such as chronic issues that happen every X time.
That's where we apply something more kind of a deep dive on the problem, which is more close to the RCA. We use for that something called eight disciplines, which is eight steps. At the end, all are based on the PDCA. Right.
And one of the steps is find the root cause analysis. But these RPS, the main value added needs to be quick in the floor right away after the issue happened and implement the actions as fast as possible. So are you effectively saying you're trading a little bit of maybe accuracy for speed? That is correct.
Yes. Sometimes we spend too much time in a not complex issue and we pretend to make it perfect. And we end up finding another issue. We forget the one that we was investigating and we didn't finish none of them.
So you would rather. Yeah. How do you kind of walk someone through the logic of that? Because I think a lot of times, again, when we see people do RCAs, there does seem to be a over sometimes like a little bit too much emphasis on it's got to be perfect.
And I feel like sometimes that the perfect is the enemy of the good enough, because in that same time of making one perfect, you might have been able to do five, you know, five different investigations, which would have yielded better results than just putting it all into that one. Yes, but the difference that I can notice is the RPS is more focused on equipment breakdowns. A pump that was failed, a bearing that was seized, a leak that happened in the line. The RCA or the depth that we use here is more related with the process.
So it's an effect that is happening with X frequency. That is not easy to identify what is the cause of the effect of the problem. That is where we apply more of deep investigation. So give an example of one of those, for example.
Yes, for the deep investigation, I can share one that happened time ago that was a difficult one. It was a defect in an aluminum chamber for a sterilizer that was cosmetic. But that equipment price was above $100,000. So no one to buy $100,000 with a spot inside of the chamber.
And that happened after a process that was passivation. So before passivation, it was not noticeable in effect. So we went to different, through all the process for the root cause analysis, make the key on the root cause analysis with the eight steps is that we have to prove that the factor is directly related with the outcome. In this case, we define the factors and we have to prove kind of a scientific method, make an experiment or duplicate the issue.
And if that factor is direct, that's how we connected with the problem. In this case, the problem was overheat. That chamber was welded with a robot and the aluminum welded is not stable. So it requires rework.
The robot welding is a very stable parameter, but the reward was not. So the operator that does the reward kind of increased current on the process and that great overheat that it was noticeable until the chamber was on the finishing process. And actually we did an investigation with a university with an electronic microscope when they see the difference on the structure and they said this is caused by heat So that kind of an extreme example That great Example No I really appreciate that walking through it It's very helpful.
Well, we can kind of keep talking about the - so you would recommend to others out there effectively, hey, let's create two tracks. A track that is rapid, still as deep as you can go, but just time box it. Don't give yourself more than five days, do it right away, et cetera. If you think if you put those parameters in place, that kind of is like the immediacy of that and the quickness of that offsets for any kind of deepness that you might be missing.
But then there's also this category for, hey, there are some processor chronic problems that we do want to do a little more deep dive into. Yes, that is correct. I would think that 80% of the issues will be solved through the basics, such as the RPS. The other 20, it requires more deep investigation, such that or the mic depends on the type of issue.
That's good. Well, one other thing I want to ask about before kind of we close, just around innovation. What are you seeing that's the most innovative in your work these days, whether AI or, you know, IoT or what are you seeing out there? Yes.
Well, AI is on all the, we are using AI actually in condition monitoring with a company called Traction. So we installed sensors for vibration monitoring. And the way that they use AI is these devices are connected on their network, and they can compare the performance with other similar equipment that they have on the network. And based on that, they kind of fine tuning the performance to detect an anomaly.
So we just finished this pilot and we are in the process to implement it on the rest of the equipment. AI is kind of today's better tool that everybody wants to use. Right. Well, and I think we're getting close to the point where is AI an innovation anymore?
It's just so built into what we're doing, right? It's like, yeah, of course, that's a thing. But it is still helpful to hear. Well, it's been great talking to you, Juan.
I always like to ask people for kind of one final word, just based on your history, your work history, you know, kind of put in your mind. You're speaking to either a younger engineer or, you know, somebody in the field. What would you want to share with them as a way to help focus them? Or even talking to yourself as a, you know, looking back on your career, what would you want to have said when you were younger?
Yes. Well, something I would say is go back to the basics. We received a visit a few weeks or a few months ago from a person, a consultant, that gave us a concept that is called FLAT. He said that 80% of the issues is caused by FLAT, which is fastener, lubrication, alignment, and balance.
So before we start putting sensors on the equipment, make sure that it is balanced, aligned, keep good lubrication. And after that, everything is good. We can go for more advanced. I think that's a good tip.
Yeah, make sure you got the basics right, the fundamentals. And once that's happening, think about adding on. And in some ways, it goes with your rapid approach to problem solving, right? Like, hey, if you're not doing the basics quickly, it doesn't matter if you're doing more because you're going to miss some things.
I think that's such a helpful insight. Well, Juan, thank you so much for being with us again. Juan's the Director of Manufacturing and Engineering at Frenesius. And so we are just so thankful to have you on the show.
We appreciate you joining us. Thanks, Sebastian. And to all listeners out there, really, anyone, we're about helping engineers solve problems. We want them to take more steps using technology where possible to solve problems.
And we hope this conversation has been helpful in that. Again, if you're looking for anything related to RCA, whether training or research and materials, check out reliability.com. And of course, if you're looking to implement RCA solutions quickly, check out our software, Easy RCA.
That's all we have for today, folks. Thanks for joining us. And thank you for being on this journey with us as we seek to make the world a more reliable place. We'll see you on the next episode.
Take care.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.