The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Leadership/The Digital Leader Show
The Digital Leader Show artwork

IBM’s AI Bet: Rebuilding Software Development with IBM Bob

The Digital Leader Show · 2026-05-07 · 39 min

0:00--:--

Key moments - from our scoring

Substance score

56 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality10 / 20
Guest Caliber14 / 20
Specificity & Evidence12 / 20
Conversational Craft9 / 20

IBM Bob represents a deliberate rethinking of how to deploy AI in software development. Unlike the industry's rush to throw massive models at every problem, Neil Sundares has architected a system that layers model selection, user experience, and guardrails to deliver consistent developer productivity gains. The breakthrough wasn't model sophistication alone - early versions achieved 30-55% accuracy, yet developers found them transformative because the system integrated seamlessly into their workflow. IBM deployed Bob internally to 80,000 developers before external release, systematically onboarding users and letting Bob itself generate 40% of Bob's codebase. What emerged surprised the team: productivity varies dramatically by role. Junior engineers treat Bob as a senior architect guiding them through compliance requirements like FedRamp certification. Senior engineers use it as a code generation workhorse. Consulting teams prototyping see 80-85% gains while teams handling complex infrastructure see lower but still meaningful improvements. The episode covers the trifecta Sundares emphasizes - model quality, user experience, and cost optimization - and how orchestrating between small and large models beats pursuing raw capability scaling.

Key takeaways

  • →IBM Bob achieved 45% average productivity gains by combining model selection, UX design, and guardrails rather than relying on massive models alone, with 40% of Bob's code generated by Bob itself.
  • →Different developer roles experience dramatically different productivity multipliers - junior engineers use Bob as a senior architect, while senior engineers deploy it as a junior developer, enabling both upskilling and acceleration.
  • →Model size alone doesn't drive adoption; seamless integration into the development workflow (frictionless recommendations that feel like the developer wrote the code) matters more than raw accuracy or capability.
  • →The approach orchestrates between small and large language models based on task complexity and cost, rejecting the assumption that bigger models always deliver better value.
  • →Programming is increasingly human language, not syntax, which expands the surface area of potential software engineers while transforming existing roles like financial analysts and operations staff into builders.

Guests

Neil Sundares

Topics in this episode

CursorSmall language modelsLarge language modelsTransformer modelsGitHubStack OverflowIBM BobFedRamp complianceLSTMscode completion

Questions this episode answers

How did IBM Bob achieve 45% productivity gains with only 30-55% model accuracy in early versions?

Early accuracy percentages misled because developers work with unreliable technology routinely; they found 30% code acceptance rate amazing because the system provided frictionless, integrated recommendations in context rather than standalone predictions. The trifecta of model quality, user experience design, and deployment architecture - not accuracy alone - drove adoption and productivity.

Why did IBM deploy Bob internally to 80,000 users before releasing it externally?

Internal deployment allowed systematic feedback loops and feature refinement. IBM started with 10 developers, then scaled incrementally to 100, 1,000, 2,000, and beyond. The team prioritized onboarding junior and mid-level engineers first because Bob is an engineering tool, not a prestige feature, and their feedback shaped the product.

What surprised IBM most about how developers actually used Bob?

Developers used Bob completely differently by role and domain. Junior engineers treated it as a senior architect guiding them through Fed Ramp and compliance processes, while senior engineers used it as a junior developer generating boilerplate code. Consulting teams building prototypes saw 80-85% productivity gains while teams handling complex infrastructure saw lower but meaningful improvements.

How does IBM Bob handle the unreliability problem with large language models?

Rather than solving unreliability, Bob architectures around it through orchestration: small models handle simple tasks faster and cheaper with known limitations, large models handle complex reasoning, and guardrails built into the user experience ensure human-in-the-loop validation. Cost optimization and speed also matter because developers need frictionless, fast recommendations during active coding.

Why did Neil Sundares leave Microsoft and eBay to join IBM?

Sundares seeks undefined, poorly-defined problems where he can shift organizational momentum and culture. After 20 years building greenfield products in high-velocity environments, he wanted the double challenge of building new AI capabilities from scratch inside a 114-year-old company, requiring him to bend rules and demonstrate incremental benefits to earn freedom.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The guest section contains genuine substance - internal deployment mechanics, the junior/senior inversion use-case, productivity variance by team type - but roughly the first 12 minutes is weather chat, tech headlines, and surface commentary that produces zero operator insight. Density is moderate, not exceptional.

your junior engineers who are using Bob as a kind of a senior architect who can help them guide through the process of Fed ramping. Then we have senior engineers who doing Fed ramp high or they're doing stuff they really need a quote unquote a code monkey
there were teams who were productive very very differently like there are for example our consulting organization where they do prototypes a lot of prototypes they were product their productive numbers are like 80 85%. Then you got people doing complex things like you know open telemetry or fed ramp etc. Their numbers are like 30 40 45%

Originality

10 / 20

A few genuinely fresh angles emerge - the Bobcoins feedback-gating mechanic, the junior/senior inversion dynamic, and the early code-completion history where 30% acceptance was celebrated - but most framing (small vs. large models, AI-as-search-engine misuse, jobs debate) is well-worn territory.

we had internal currency called Bobcoins um which translates to some dollars and we would just say hey you get 50 bucks worth of Bobcoins and then when you come to that 50 bucks you got to write a feedback
a machine being wrong two out of three times and still the developer saying wow this is amazing

Guest Caliber

14 / 20

Neil Sundares is a genuine practitioner - GM at IBM after VP of Engineering & AI at Microsoft and head of eBay Data Labs - who credibly traces a personal arc from early IntelliSense-era code-completion to running an 80,000-user internal deployment. This is real operator experience at scale, not thought-leadership.

we are today at 80,000
when I first talked to uh our CEO there uh about bringing AI machine learning and developer productivity together. It's like the question was where do you come from? How do you bring AI and compiler systems together?

Specificity & Evidence

12 / 20

Concrete numbers are present and credible - productivity ranges by team type, sprint story-point increases, 90-95% weekly return rate, the $50 Bobcoin allocation - but key claims like the 45% average productivity gain are introduced by the host without methodology, and the 'Chroma research' citation is too vague to verify.

they used to do seven eight story points now they're doing 12 13
our consulting organization where they do prototypes a lot of prototypes they were product their productive numbers are like 80 85%. Then you got people doing complex things like you know open telemetry or fed ramp etc. Their numbers are like 30 40 45%

Conversational Craft

9 / 20

A handful of questions are genuinely good - 'what surprised you most about what developers actually did with it' and Chris's 'unreliably unreliable' framing - but the hosts never challenge the productivity numbers methodologically, let the guest run long without redirecting, and burn ~12 minutes on banter and headlines before the interview begins.

I'm curious what what surprised you most about uh what the developers and the users actually did with it
they're unreliably unreliable and that's very difficult to engineer around

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

build25software20back20models20code19human15experience14today13first13question13system12language11microsoft11neil11engineer11machine10

Episode notes

On this episode of The Digital Leader Show we are discussing IBM’s entirely new AI-first development environment, IBM Bob…PLUS, headlines for digital leaders from the world of enterprise technology. About the Show: During each episode we explore some of the key issues defining our ongoing transformation into digital enterprises, including automation, AI & digital transformation. Join us for a unique, unfiltered discussion at the intersection of technology and business, helping you to become a better Digital Leader. About the speakers: This episode’s guest is Neel Sundaresan, General Manager of Automation and AI at IBM. Neel has previously served as the Sr. Director & Head of eBay Data Labs and as the VP of Engineering & AI at Microsoft. Chris Surdak is a Senior IRPA AI Advisor & Analyst and was formerly White House Chief Transformation officer, Automation & AI Practice Lead at EY & Executive Partner for Digital Transformation at Gartner. He’s an engineer, futurist, transformation executive and best-selling author, with over 30 years’ experience in technology development and deployment, digital transformation, blockchain, data and analytics and AI & intelligent automation.

Full transcript

39 min

Transcribed and scored by The B2B Podcast Index.

Kind: captions Language: en Ladies and gentlemen, welcome to the digital leader show where we talk key issues and news at the intersection of business and technology in hopes of making you a better digital leader. Hosted by Dan Goodstein and Herpa AI, a professional association, media, and peer research network providing education, networking, and advice to over 130,000 global members on automation, enterprise AI, and digital transformation. We're live Wednesdays at noon Eastern, 9:00 a.

m. Pacific on LinkedIn and YouTube, and the audio recording is available wherever you get your podcasts. And now, here's your host, Dan Goodstein. Hello everyone, welcome to the digital leader show.

Today we're discussing IBM's AI bet, rebuilding software development with IBM Bob, plus headlines for digital leaders from the world of enterprise technology. My co-host is Chris Serdak. He's an AAI senior analyst, award-winning author, and former White House chief transformation office adviser. He's also led teams at EY Gartner, uh, and a bunch of other places.

Chris, man, it's been a while. It's been a minute. How you doing, man? >> I've I've missed you.

I was saying to one of my friends the other day, it's hard to believe the year is almost half over. >> You know, I I every year I plan for that. And every year I'm still surprised by that. But uh yeah, you know, if you live uh on the eastern seabboard uh here like I am in New York, you know, we we we've lost spring, so it just goes from like winter to summer.

And so you just kind of forget what uh you lose your orientation there. Well, I for for people like in the East Coast, I'm going to move I'm going to move out to California because it's nicer weather. It's been raining on and off for the last several days. And this morning during one of my Zoom calls, I had an earthquake.

So, you know, there you go. >> Oh, geez. Yeah. >> Pick your poison for natural >> casualty of California.

Yeah, exactly. Exactly. Well, it's good to see you, buddy. We got a great show today.

a really interesting guest. So, I don't want to uh uh delay too much, but you know, a couple couple really interesting things going on in uh as far as headlines go. Um the one that came out uh recently that kind of got my attention is Meta's $2 billion AI deal shut down by the the Chinese uh authorities. This was an acquisition of uh Manis, an Aentic AI startup.

And um this is according to Bloomberg really really um a uh a turn for for Meta who's kind of really counting on this. Um and then you see more and more of this you know regulatory trade war type of stuff. Um makes you wonder how do you really you know innovate and plan for for these global economy type of situations. >> I you you stole the word out of my mouth.

war. I this to a degree this is intellectual property and economic warfare. Um and I think we can expect to see a lot more of that to to your point about meta use you kind of hoping to acquire its way back into uh at least competitiveness if not a lead in the AI field. Um I think it's definitely an interesting shot across the bow.

Um but it also there's been this long time perception that you know China's way behind the US companies and so forth and I think deepseek would would argue uh otherwise but it is interesting that uh you know we have companies that are going to Chinese startups and trying to buy buy that IP from them and and China itself is stepping in. Uh the FCC uh in the US is also expanding its foreignade Wi-Fi router ban uh and it's going to include portable hotspots, cellular home internet devices.

So um I don't think they explicitly mention China in that uh document, but uh you can imagine where a lot of that stuff is made, right? >> Yeah, I've I've been talking for over a decade about how a lot of this, you know, these exploits and so forth um are built into the hardware if not like actually etched on the chips. So, you know, again, nothing surprising there and and um the the value of information is such that people are going to do whatever they can to get their hands on it.

>> Yeah. Yep. Speaking of um lots of news as always on uh Open AI, the uh Musk versus Open AI trial starting. Um so that you know is all around the the you know who's going to lead it and and the nonprofit status change and uh all that.

Meanwhile, Microsoft and OpenAI just redrew their partnership uh extended it to 2032, changed uh removed this kind of controversial AGI clause and um uh does not no longer limits uh Open AI to uh to uh Azure, although it'll be kind of an Azure first uh plan seems like. >> Yeah, I I I saw Musk kind of going bonkers on on OpenAI and Altman on on X. So, um, I fully expect a a court-ordered gag order shot across the bow within the next probably 24 hours. Um, you know, not bad mouthing the counterparty and on social media.

Um, and then that whole AGI thing, I I don't know how you pull that off, although I don't know how you pulled it off to begin with, but then to pull it out of the clauses. Um, you know, I'm not a practicing attorney, but I did go to law school, and this we're rewriting entire chapters of American law uh as as we as we go. It's absolutely fascinating to watch. Uh and the last piece about open AI is uh some reports came out this past week that they're plotting a smartphone uh of their own to pair with its, you know, big AI, you know, AI everywhere, AI in the home uh kind of kind of uh push.

You know, I I I can't underestimate them. You know, that would be stupid. But um my first reaction was again like, do we need another smartphone manufacturer? I feel like we've seen this movie before.

Uh, and ironically, one of the companies they're talking to is Qualcomm, who, you know, really, you know, struggled, I think, uh, in the last round of cell phone wars, right? >> Well, as as you just said it, I mean, do we all remember the last time Microsoft tried to do a phone? >> No, we don't. >> But they did try, right?

>> I remember that. I'm old enough to remember that. >> I think I I think I might have had one of those phones for about three weeks or something. Um, so >> Black Blackberries, Microsoft phones, and uh >> I was I was talking to my my team earlier this week about remember Nexttel pushto talks.

>> Oh, I love those. >> Right. So, everything old is new again. But, you know, all right, open eye, you want to get into the hardware game, too.

Fair play. But, it didn't work out so great last time Microsoft tried to do that with with mobile devices. We'll see. >> Yeah.

Yeah. The last kind of thing I wanted to hit on before we uh invite Neil, our guest, to join us is, you know, and I know this is a topic you've talked a lot about, um, but, you know, there's this there's this article that came out in the Wall Street Journal, um, talking about whether or not the the tech giants are going to regret all of their layoffs, you know, and and and and their their push to go all AI. Um, Meta is reportedly uh laying off about 10% of its workforce to offset the soaring AI infrastructure uh costs.

And um uh Microsoft just offered its first ever voluntary buyouts. Uh the list goes on and on. But um the the interesting part, not not surprising to you because I know you've you've rang this bell before, but um you know there's some really interesting reports coming out now that researchers coming out that's saying that uh AI compute is now more expensive than the humans it was supposed to replace. >> You can't make this up.

Although, you know, we all do in Excel spreadsheets that we call ROI calculations. We make up stuff all the time. Um yeah well and we've you and I go way back with like outsourcing and RPA and so forth. We've been trying to convert capex into opex for a quarter of a century and now we're trying to convert opex back into capex.

It's all funny money. Um, so I the um I think what we're also seeing I I have about 400 pages of research papers I have to read over the course of the next week just and a lot of it is on this where they're finding okay you kind of replace people with the AI but the people that are using the AI now are getting dumber and like quantifiably dumber quantifiably worse results um and a lot of the over yeah overpromise and underdel how many people are clawing back yeah people that they laid off thinking I don't need them anymore when they realize okay the the AI is replicating what they did but they don't know what they did.

So um you know I think we're going to be uh going through a tremendous cycle of of HR related workforce related angst. It's actually onethird of my upcoming book on this stuff kind of talking through my predictions of what that'll look like. And I don't think there's any easy answer. Um, other than we're going to have to go through it and as you say, I think a lot of you're going to have to hire back 5% of the people that you let go and you're going to have to pay them double what you what you used to if if if they're going to come on board.

So, um, every everybody wants something for nothing and the second law of thermodynamics remains undefeated like I say all the time. >> Interesting. Well, that's a good segue. So, uh, let's let's introduce, uh, our our guest for today.

It's Neil Sundares. He's the general manager of automation and AI at IBM. He's previously served as head of eBay data labs and the VP of engineering and AI at Microsoft. Neil, thanks for joining us.

>> Thank you. Thank you, Chris, and thank you, Dan. Uh, appreciate your welcome. Yeah, thank you.

>> Pleasure to have you. >> Pleasure to have you. So, um, I guess the obvious question, Neil, I wanted to ask to start with. I know Chris probably has more more insightful technical questions than I do, but you spent 20 years at Microsoft and eBay.

Uh, you could have very easily done the work that you're doing now, uh, at IBM there. Why why did you make the jump? Why did you believe IBM was a better place to kind of build something new? >> Yeah.

Uh, it's a good question. I think u you know in Silicon Valley or even in Washington where I was with Microsoft I think every few years you look at you know what you're doing and what more can you do and do you have an environment to do that or not and almost in my entire career I've made career jumps asking that question because there is a part a time in your career at this particular company where you ask that question and then also you know at this level people reach out to you they want to hire you etc and you have lots and lots of conversations and you say, "Oh, I this is my next idea.

Can I do it in this place or not? Do I have the freedom to do it? Do I have the facility to do it? Um, do they understand me?"

And you ask these questions and you discover the answers for yourself and then you make the move. And that was my move from eBay to Micros I mean before eBay I did a startup from the startup to eBay from eBay to Microsoft and Microsoft to IBM. while the role defines something. Now, I'm a sucker for challenges.

I'm a sucker for, you know, undefined problems or poorly defined problems. So, moving from the Bay Area to Washington was not easy. Um, you know, because here you walk up to a coffee shop and hire 20 people. Uh, you know, it is not the same thing there.

Uh, you're trying to teach data to somebody who basically built products with features and never looked at data. So in some sense you take on those challenges and you say well can I make a difference here right um and you know do you need to do that no do you need to take those risks no so there are different kinds of risk you know one kind of risk is going to a startup and taking the risk another kind of risk is going to a big company and saying can I go change the momentum here change the culture here or do what I want to do here so I always say you know can you be the rattlesnakes in the belly of the elephant um and and and you know put out Right.

>> Well, and and you you you me you kind of allude to, but it to me it sounds like a double challenge, right? One one is is, you know, shifting, but the other is you're basically trying to build this from scratch inside a 114 year old company rather than kind of bolting AI into onto existing tools. And I think IBM kind of historically has has done it the other way, right? >> Yeah.

And and I think you know credit to my my my peers my managers uh who let me do that right so my my approach to doing these things is I never ask for oh I need to do this I need to do this new initiative give me 500 people I don't know what to do with 500 people almost everywhere I've done new initiatives or new new products or new new innovations I always say I don't know what to do with a lot of people so let me pick the people I want and then we go from there and this is disturbing to a lot of you know executives who just say oh how serious is this guy right and then you play the way you want to play right then you break the rules you want to break then you challenge the system the way you want to challenge and tell them and and then show them incremental benefits of that till they kind of understand you right so when I first went to Microsoft I was in a team meeting and I I showed all kinds of data and they were like oh we never look at data this way we just build features right when I first talked to uh our CEO there uh about bringing AI machine learning and developer productivity together.

It's like the question was where do you come from? How do you bring AI and compiler systems together? Because what I was pitching was we have tons and tons of code data um you know at that time it was Stack Overflow and GitHub etc. And there's enough information in there to automatically recommend things to the user.

At that time I mean the models were not great. In fact, the first couple of models we put out were, you know, traditional machine learning models, but there was a problem to be solved, right? We just, in fact, the first, if you look at go back and look at even precot, look at Intellico, it was like developers writing code, they put a dot and they want to make a function call, an API call. Can I make the right API call there?

Can I recommend the right parameters to the API call? And I've done machine learning for many years and typically the numbers would come at 90% 95% accuracy but um you know here we were getting 50 55% accuracy acceptance rate was 40%. So that would in any machine learning uh scenario or you know evaluation system would be poor but developers came and said wow this is amazing how do you do that and then we went into they wanted more and more. We went into deep learning.

We went to LSTMs and then the transformers came along. We built our own transformer models. The first models with just few hundred million parameters but all we were doing is code completion and the code acceptance rate was you know 30%. And that was great for developers.

I mean you imagine a machine being wrong two out of three times and still the developer saying wow this is amazing. Right? So that's how it started and that's music to your ears. Um and then you go from there and now we are here.

Right? Uh it's not about lines of code anymore. It's about you know how much can be automated and how much can be augmented. Yeah, >> if I may.

It's it's funny you say that because even um we came across this a lot when I was building spacecraft. So as you were just saying you know 30% how is that good? Engineers work with unreliable technology all the time. I think one of our challenges with LLMs right now is that they're unreliably unreliable and that's very difficult to engineer around.

Is that kind of your experience as well? Yeah, very good point. So, how do you build a guidance system when you have I mean, of course, as as the models have become more and more complex and more and more capable, they don't tell you where they're capable and where they're not capable, which is kind of your your way of saying they're unreliably unreliable. So, what guardrails do you put in the system?

What kind of framework do you put so that the users accept your the recommendations in a way that actually human intelligence kind of mixes with machine intelligence and kind of uh you know puts down the machine stupidity or you know the unreliability right simple thing I would so when I built even the first version of this thing it was sort of like it is not just the models it's a model it's experience and it it is also the the architecture in which you deploy because I could have fancy model which takes two minutes to come back and answer a simple question or it could come back super fast but I could build a really bad user experience because think about it engineering software engineering is an analytical experiment analytical experience so people who are we're not shopping we're actually writing code we're doing math we're doing and in the in the middle of that the computer comes and makes a recommendation right it needs to be smooth it needs to be frictionless it also needs to make make me it make it look like I wrote the code not the machine wrote the code.

So how do you build that experience into this thing? So I always say it is a trifecta of the model the quality it produces the experience you provide and in fact today with the models costing so much it's also the cost of the model now do I really need a sonet 47 to do some simple task or can I take a simple model and then put it in because I have the advantage of speed small model and I also know that there is human being in there. So how do we orchestrate between these large and small models in a way we can provide the best experience to the user is kind of what what we always think about.

So it's not just cost-saving it is not just performance it is it is a combination of all of those things combined with experience. Yeah, I think you touch you touched on that. A lot of people are going back to the or at least considering small language models and I don't similar to you worked in machine learning and analytics for a very long time and so well I have a 100 million parameter LLM like you said and then all right let's do a 100red billion one well if you're asking the question what color is a strawberry how you know how much more red is it going to be right once you have the answer red is is the is the is the system that costs a hundred times as much going to be redder >> no so so you know if if I'm scale ing up what are the different questions that I'm asking so I can get different insights.

>> Yeah. >> Yeah. Exactly. And if you look at today, right, if you look at today, >> the simplest task is code completion.

And if you look at the workloads, you know, probably maybe 1% 2% of the workload is really code completion. Most of the time I mean we're abstracting away in some sense programming is human language. It's English. It's Chinese.

It's Japanese. It is not Java. It's Python, C++. Right?

And that's why on the one hand we say oh the software engineering jobs are going to go away. On the other hand it's really more and more people are becoming software engineers because now my financial analyst can be a software engineer. My operations person my HR person can be soft engineer because they all speak human language right. So we are actually increasing the the the the surface area of people who can build software.

So it's kind of an optimistic way of thinking about this thing. Right? So um I think we are in very very good times with AI. I mean I know I heard a little bit of what you said before.

Um are we not going to have software engineers anymore? We call everybody software engineer. How about that? Right.

Speaking of Neil, I mean you you guys deployed this internally first, right? I think 30,000 uh internal IBM users, developers, and achieved something like 45% average productivity gains. I'm curious what what surprised you most about uh what the developers and the users actually did with it. >> Mhm.

>> I So, you know, we are today at 80,000. So, just to be uh the latest number is 80,000. So but we were very systematic in the way we would deploy it because I started with a small team which is I mean you know you you already talked about IBM this is not typical in IBM right and and I said well we are building a productivity tool which means Bob's a productivity tool the Bob team should use Bob so how much of Bob can be generated by Bob right so features in Bob was written by Bob so literally I can kind of do a backup calculation say 40% of Bob is written by Right.

And I had the first 10 developers within the team and a little bit of the periphery use it. Then I had 100. So I had about 2,000 people in my organization. Then I had 100 people use it.

Then I have 200 people use it. As they use it, they kind of give feedback. They give feedback. You know, you grow the you grow the software.

You build in features. Some features are important, some features are not important. So just getting bootstrapping that way and growing it. Then we went to thousand users, 2,000, 5,000, 10,000, 20,000.

And I was we were very particular about who we on board first. So it was really because we were thinking of it as an engineering tool. So people would come and say oh I'm a distinguished engineer. I'm a fellow or I'm a VP but I could.

So sorry you get in line. You know that software engineer band seven band 8 in IBM which is kind of junior software engineer staff software engineer. They need to come and use it. And what was surprising of course was you know we we didn't even have a manual.

Um and of course people are used to other tools like cursor etc. So it's not it's not like they're they're coming completely blind into this thing but you know without any manual they they would start they would make mistakes and iterate and very quickly you know use the best of the features available in the system and because most of the time they were actually programming in human language most of the time it's English but we also have Japanese Korean etc users programming in those languages and then all we needed to do was build belts and whistles like oh there is this mode, there is a code mode, there's an architect mode, there is a designer mode, there is an ask mode and then we need to most of the time just change the interfaces around and of course we are also looking at user experience.

We also want to know we do offline eval to say which model works really really well. So if I were to summarize on what surprised me like how different people used it differently because mo the question that was asked was how productive people are and the answer is different people are productive very very differently right so if I'm a junior engineer let's say you know in IBM we require all of our products to be fed ramped and typically before that fed ramp developers need to be senior they need to know what it takes to make their make the software um you know ironclad so that we can deploy on fed environment.

Now your junior engineers who are using Bob as a kind of a senior architect who can help them guide through the process of Fed ramping. Then we have senior engineers who doing Fed ramp high or they're doing stuff they really need a quote unquote a code monkey and they can use Bob as a junior developer to do that. So you kind of see the both ends of it. The other thing that was surprising was you know people ask for lines of code or productivity etc.

And then what we found was um there were teams who were productive very very differently like there are for example our consulting organization where they do prototypes a lot of prototypes they were product their productive numbers are like 80 85%. Then you got people doing complex things like you know open telemetry or fed ramp etc. Their numbers are like 30 40 45%. Uh we also I also saw that people were doing more story points per sprint than they did before.

you know they used to do seven eight story points now they're doing 12 13 so in some sense incrementally they're getting bolder and bolder in the way uh they do their story point so a lot of these were surprising positive and also I thought that maybe we'll get 50% usage maybe people will do their job and be done I mean like there is this thing about John's paradox right you know if things get cheaper you'll use more of it there are positive and negative sides to it and what you see is that people were engaged more in the weekends people there was a guy who was taking his laptop uh to his daughter's ice hockey game.

Um you know stuff like that. So people were more engaged and the story you hear is I would never have touched this because it was a boring job or it was a difficult task etc. Now I can actually even attempt it and do it and I feel comfortable doing it. Um know there was an engineer who came back from maternity leave and she was like I came back and I had no idea what uh you know I I felt like I lost six months of my life.

Um and now I I when I go and ask other engineers question I feel really really stupid. I got access to Bob now actually I don't have to look stupid. I feel like okay I want to continue with my job because I have my partner here. So those are all somewhat emotional but at the same time very strong stories that we heard.

So the engagement was pretty good like you know we have 90 95% of the users come back every week and we did put some systems in place for example we didn't just give away everybody free blank check go and use as much as possible so we give everybody I mean now we give more but you know we had internal currency called Bobcoins um which translates to some dollars and we would just say hey you get 50 bucks worth of Bobcoins and then when you come to that 50 bucks you got to write a feedback you would say how you used it What did you find useful?

What would you not find useful? Have you used competitive tool here? How did it help? Uh and then we ask for how productive were you and explain that productivity.

We also ask for happiness score. Did it make you h was it a delightful experience out of five? Because you know we engineers are grumpy people and keeping them happy is an important thing. And so we got amazing I mean like amount of data from the users how they are using it.

we could put it back into the system. We built an entire slack channel for the developers and Bob would I mean we also automated a bunch of things for Bob to answer some of the you know mundane over and over repeated question how do I increase my Bob coins or how do I connect to a MCP server stuff like that but we also got those feature requests and bug reports automatically you know created back um um u backlist for our developers to work on assign it to developers and all so that's why with a small team we could actually do a lot of stuff that cannot be done with large team so It's a total productivity tool for the Bob team itself, but also for the teams that were using Bob.

So, a lot of a lot of the enterprises I speak to, Neil, are are struggling to move from kind of AI experimentation to to real productivity. What's what's your sense as to what they're probably getting wrong? >> I So, there's one thing that I mean, I would say there are two or three things that I can talk about here. So you know I've talked to several customers.

We have a ton of enterprise customers. So typically what happens in a CIO organization is some engineer or set of engineers come and say hey get us some GPUs and we can build our own you know LLMs uh that will make us productive. So that's one sometimes and then they don't they don't get too far because you know it's not very easy to build you know LLM models that will work for you and always you know the frontier models or or people who are whose whose full-time job is to build models will be much better than you or sometimes they say hey get us the models and we'll build applications for you that will make us productive even that doesn't work because it requires you to know how to build these you know idees or experiences on top of the models So when you hear numbers like 95% of the you know projects failed it is because they have made the wrong investment in these things in in my opinion and the third thing is the users themselves many users kind of use you know AI assistants as search engines they just ask questions and hope that it works and and then accept it as is and then they they the results they accept are not great and then they have to go back and fix it and they lose trust and they use it less and there's this great work from chroma research we talk about context rotate study where they studied 18 models and sees how performant degrades and performance degrades so in some sense you can I don't have scientific proof but at the same time you can see as the experience it's almost like a lemon market as the experience become gets worse and worse they get discouraged and they use it less and less and less so building an AI tool AI assistant or or call it coding assistant whatever that might be is a lot more work than just models right and using it is also a lot more work than just writing code right it is you know there's a ton of work to be done in context engineering and as to tool builders like you need to have the entire information ecosystem you need to understand multi-turn stateless stateful workflows uh you know you need to curate what the model sees and you need to put in the right context etc for example if If I were to build something that understands 10 PDF document, I can't just pl the 10 PDF documents onto it and lose the context.

I have to arrange it in some in such a way that it is indexed appropriately served through an MCP server. So there's a whole bunch of stuff that has to be done. And of course these are manual tasks you would do if you don't have the right tool. So we as tool builders, we have to provide you with the right tools, right experiences so that we can automate as much of that possible more people get it right.

Right. I think we are kind of sort of in the middle of that journey at this point. >> Yeah, Neil, I I think that's such a critical point that you make that so many people are using AI as as a super fancy search engine. And uh you I I say all the time using probabilistic tools to solve deterministic problems deterministically causes probabilistic failures.

>> Exactly. >> The worst possible use case for the tool, >> right? >> So, you're laughing. Does that actually make sense to you?

Because Chris has been saying that for months. I have no idea what he's saying. What the heck? >> No, it makes sense.

I mean, it is it is really and I've never said it myself this way, but it makes total sense, right? It is. >> Well, so so how do you how do you kind of encourage people to get out of that, hey, it's it's Google with with better dictation and make it more, hey, this is actually something that's helping me think differently to get different results. >> Right.

Right. I mean, it it is right. So for example when we initially put out Bob internally and there was so much success other teams which are building you know there there were other products and they have they have they want to bring in AI they were like oh can you give us access to the latest uh cloud models or or you know open AAI models and they want to just slap it in and build a chat and face no that's not how you build a system you need to have the right experiences you need to have the right eval there's a discipline to building these systems and in fact I often say you know this is the best time for UX designers in AI while AI can build you interfaces but you know having human machine interaction in a seamless way in a way that it actually works very very well is critical right so I'll give you some examples so for example if I say hey Bob look at this code and tell me what is what is the problem with this is there any security violation can you help me fix fixing with it fix it one way to do it is to say here I fixed it go deploy another way to is to educate you here are four options right option number one option number two option number three option number four then Bob tells you if you use option number one this is what will happen and this might fail option number two this is a slight hack right and so on now you not only are using AI to do your job you're also getting you might pick one of the four options or you might even ask Bob to say hey pick one the best option but in the process you're also getting educated about the other options Right.

So it's kind of I mean when you build AI as an educational experience while being productive it's almost like a project I mean why do projectbased classrooms succeed? It's because of that, right? You you try, you ask questions, you fail, you succeed, and next time you come around, you've actually done better, right? So building that experience is an art.

I mean, it's it's not like you run you let clock code run for three days and some magic happens and then you don't know why it worked or why it did not work, right? Maybe we'll get there when we understand all of the possible task. But I think at this point it is building these experiences is really really important especially when you're taking on I mean we looking at the prompts from say back in June to today there are people are throwing more and more things at AI which means the responsibility of people who build build tools is higher and higher.

Neil what do you what do you think uh where do you think this AIdriven development is going to head over the next two to three years? And if anybody's watching any digital leaders, CIOS, CTO's and they're they're evaluating this kind of thing, what should they be preparing for? >> I mean based upon at least what I have seen and what I've experienced, the way we develop software is changing, right? So is it is it something that we would have imagined five years ago when we had just line completion and models that were wrong most of the time and you know they they kind of had frozen training data etc.

um no but today they are good enough and we can build experiences for example if you look at our workloads right 60% of our workloads are modernization workloads they're not like brand new project that you're working so if you look at a software development life cycle it's not just writing code compiling building and deploying and scaling it is also taking existing system optimizing them finding non-functional issues and modernizing them migrating them to new languages or new frameworks like take for example Java you could have monolithic servers written in Java 7, you might want to migrate to Java 21 and you know containerized uh systems etc.

So you know and so there's a new way we build it. And the second thing is that all of these have become now agents. Now there are some deterministic and some probabilistic. How do you build a software system that's a combination of agents and these agents not only talk to the human, they also talk to each other.

Now today they talk in human languages. I mean I think to Chris's point you know they used to talk APIs now they speak English or human language now APIs were very specific so every time you want to have a slightly different interface you needed to change the API now all of while all the underlying information was there today what we do is build minimal APIs and make rest of it part of the conversation so there is a deterministic part and then there is a probabilistic part now the probabilistic part you have to get right right so you have to say Well, you know, the question is going to be asked in English, let's say, or Japanese, and the answer is going to be given in Japanese or English.

But the underlying system is deterministic. How do I make it work? As we go even forward, agents are going to talk to agents and make decisions. Is this answer right or not?

Or do I need to have a back and forth negotiation? Right? All of that you have to get running very well to the extent they to to the amount they ran well before when we had just doing software testing, right? and they al I mean they got and also you need to optimize on how many back and forths you have because the systems have to run at scale with efficiency etc.

So that's kind of where we are entering the software era. So if I'm a college student today and I'm going I'm studying computer science, I'm going to enter the software field. I should be trained very differently than I was trained. There are some foundational principles in software development that don't go away.

Right? In fact, if you look at even human language coding, Donald Smith came in 1985 and and said, you know, the whole idea of literate programming came from his idea except that the tools were not there. So that's why he in he invented the literate programming paradigm. But his idea was if I want to get something done, I want to talk to a human being in human language which is English and that human being is talking to a computer and getting the task done to my satisfaction.

That was the vision back in 1985. All we doing is making it real today, right? Um because we have the tools, we have the systems, we have the AI systems that are ready. I mean, not I think 85 years old now, probably very h happy with what he suffered years ago, right?

>> So, if I I think we're getting close to time and given that you just brought that up, I I have to ask I think probably the most critical question our audience is would ask you if they were in in this chair and that is this. So, Neil, do you think that through Bob IBM is finally going to let us get rid of cobalt? >> Uh, sorry, I didn't get the question again. Finally, >> is Bob going to allow us to finally get rid of cobalt?

>> Finally get rid of I don't I I don't see I think you need to get rid of cobalt, right? I mean, frankly, you shouldn't care about the underlying language, right? So, even Java migration, you think about it. The re there should be a reason for me to migrate.

In fact, you should not care about cobalt, right? There are there are some of our customers who are migrating to say Python etc. But they need to get see benefit of migrating, right? They're migrating to new platform, new framework.

They're migrating to the cloud and these languages are not available. And also you see new talent that's coming in. They don't know cobalt. They don't I mean it's not just cobalt.

They don't know PLX. They don't know PL1. They don't know many other like right other languages. But they need to be able to work with these systems that were programmed in those languages.

The only thing that you have between you and the computer is a human language. So how do I take that human language and and translate it? First I need to understand all of the you know quote unquote legacy code. Then I need to say what I want to do with it.

I want to get an architecture for the system that's based upon the thing. Then I want to I mean and that is reverse engineering a lot of code that's not very well documented. Right? Then I say okay now I want it to run efficiently in this area or that area.

Some of them will get rid of cobalt. Absolutely. Some of them will migrate from the mainframe to cloud or or new infrastructures. But others might want to stay there because there's a secret sauce there and or they're too conservative to migrate.

So, we kind of go with where the clients want to be, but we want to bring the best of the tools for them. Yeah. >> Neil, thank you so much for for sharing this story and uh would love to check in with you in I was going to say usually I say like next year, but in this world I I probably have to check in with you in a couple months, next week to see how things are uh progressing, but thanks so much for joining us. >> Yeah, great great chatting with you Dan and Chris.

Appreciate your time. Thank you. Thank you, Neil. >> Thanks everybody.

Thanks for joining us. We'll be back Wednesday at noon Eastern, 9 Pacific. In the meantime, if you need help with anything you're working on, we're here to help. Please go to herai.

comanalysts for more information. Thanks everybody. Have a great day. >> Cheers.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • VC: How Benchmark Picks AI Winners - Max 10 Bets a Year, 5 Partners | Chetan Puttagunta (GP)The GTMnow Podcast · on Cursor89 / 100
  • How do you turn AI coding chaos into a repeatable playbook?The Stack Overflow Podcast · on GitHub86 / 100
  • Ep 90: AI Pioneer Jürgen Schmidhuber on the State of AI TodayUnsupervised Learning with Jacob Effron · on Large language models85 / 100
  • The End of One Model to Rule Them All: Why Enterprise AI Is Going Small, Specialized, and Multi-ModelDisambiguation · on Small language models85 / 100
  • Small Models, Massive Wins: The New Shopify AI FormulaBeyond The Pilot: Enterprise AI in Action · on Cursor85 / 100
  • How B2B Marketers Use AI to Personalize at Scale for EnterpriseB2B Marketing with Fexingo · on Large language models82 / 100

More from The Digital Leader Show

All episodes →
  • The Age-Gated Internet & Digital Rights
  • 2025 Review and Top 5 AI, Automation & Transformation Predictions for 2026 📱
  • 2025 Review and Top 5 AI, Automation & Transformation Predictions for 2026
  • The Art of Stakeholder Whispering
  • The Art of Stakeholder Whispering 📱
Explore the best B2B Leadership podcasts →
All The Digital Leader Show episodes →