The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Data & Science with Glen Wright Colopy
Data & Science with Glen Wright Colopy artwork

Keith O’Rourke | The Logic of Statistics

Data & Science with Glen Wright Colopy · 2022-08-02 · 1h 13m

0:00--:--

Key moments - from our scoring

Substance score

47 / 100

Five dimensions, 20 points each

Insight Density10 / 20
Originality11 / 20
Guest Caliber11 / 20
Specificity & Evidence7 / 20
Conversational Craft8 / 20

Keith O'Rourke, consulting statistician and Health Canada researcher, challenges conventional approaches to statistical education and practice by grounding statistics in philosophical inquiry rather than mathematical mechanics. Drawing heavily from C.S. Peirce's work on logic and abduction, O'Rourke argues that scientific statistics is fundamentally about forming future-oriented expectations about the world through representation - specifically, treating probability models as 'fake worlds' that let us work through counterfactuals we cannot observe directly. He demonstrates how Bayesian workflow, developed by Michael Betancourt and Andrew Gelman, enables scientists to interrogate whether their models generate plausible data without requiring deep mathematical expertise. The conversation addresses a critical gap in statistics education: students learn MCMC mechanics and formulas but rarely engage with the scientific reasoning that justifies model choices, prior specification, and data generation assumptions. O'Rourke's semiotics-influenced perspective - viewing models as representations standing in for the actual world - provides a framework for teaching statistics to early-stage data scientists more directly. This episode will resonate with practitioners frustrated by the disconnect between statistical coursework and actual scientific practice, and with anyone seeking to understand why assumptions matter more than computational fluency.

Key takeaways

  • →Scientific statistics is fundamentally future-oriented, forming expectations about what will happen in the world, not merely describing the past like descriptive statistics does.
  • →Probability models function as 'fake worlds' that allow us to explore counterfactuals and work out what would repeatedly occur, then transport that understanding to real observations and compatibility assessments.
  • →Bayesian workflow using ancestral simulation - checking priors, generating fake data, comparing to observations - can teach scientific reasoning to students without heavy mathematical prerequisites before they learn MCMC sampling.
  • →The disconnect between statistical mechanics (turning the crank on formulas) and scientific reasoning means practitioners can easily operationalize computation but lack guidance on the nebulous, non-formulaic aspects of good science.
  • →C.S. Peirce's concept of abduction - coming up with hypotheses through weak but necessary inference - remains underexplored in statistics education despite being central to how scientists actually form models and representations.

Guests

Keith O'Rourke

Topics in this episode

C.S. PeirceBayesian workflowAncestral simulationKarl Friston free energy principleAndrew GelmanMichael BetancourtSandra GreenlandMarkov blanketSemiotics and theory of representationsMCMC sampling

Questions this episode answers

What is scientific statistics according to Keith O'Rourke?

Scientific statistics is the purposeful use of statistics to learn about the world in the future tense - forming expectations about what will happen when you take actions, distinct from descriptive statistics which merely describes the dead past.

How do probability models work as representations of the world?

Probability models function as 'fake worlds' or abstractions where you can learn all the implications without the constraint that you can't re-observe data; you then transport that understanding to assess how compatible your real observations are with that model.

What is the Bayesian workflow and why is it useful for teaching?

Bayesian workflow involves simulating from priors to check them, then simulating data from those priors to see if the model generates plausible observations matching your data; it teaches scientific reasoning through simulation without requiring heavy mathematics or MCMC.

Why do young statisticians struggle with connecting statistical mechanics to scientific reasoning?

Students can be trained to operate formulas and code within months, but scientific reasoning - deciding on models, choosing priors, assessing data compatibility - cannot be operationalized into a single formula or workflow, creating a gap between what can be tested and what constitutes competent science.

What is abduction in C.S. Peirce's logic and why did he struggle with it?

Abduction is forming hypotheses or educated guesses, which Peirce called a weak but necessary form of logic; he struggled to justify it rigorously but recognized that all inquiry requires it, and statisticians must grapple with hypothesis formation in model specification.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

10 / 20

The episode contains a handful of genuinely interesting ideas - probability models as 'fake worlds,' counterfactual reframing of distributions, the semiotics-statistics bridge - but they surface slowly amid long tangents about academia, publishing politics, and education philosophy that add little informational value for a practitioner.

you can think of these, um, probability models, okay, as fake worlds. Now, because it's an abstraction, you can learn everything you want about that world
the biggest predictor is where they went to graduate school. Well, that's what I learned to do in graduate school. So that's how I'm going to do your analysis

Originality

11 / 20

The Peircean semiotics frame applied to statistical modeling is genuinely uncommon, and the Friston free-energy principle connection to statistical reasoning is an unusual angle, but much of the episode recirculates ideas already established in Gelman/Greenland circles without extending them in a notably contrarian direction.

if organisms survive they've actually somehow figured out how to implement uh, scientific statistics
you started out a long time before I took statistics. I studied something called semiotics, which they call the theory of science, or better today, theory of representations

Guest Caliber

11 / 20

O'Rourke is a credible practitioner with real exposure to elite researchers (Rubin, Cox, Gelman, Greenland) and genuine applied history in clinical meta-analysis and regulatory work, but he is not a widely recognised name and the conversation does not surface work done at significant scale or impact.

When I first started in statistics back in the mid-80s, I was at, uh, the University of Toronto, and I started working with some clinical researchers and I worked out for them how to do meta analysis
Don Rubin told me is that whenever you're learning a new area in statistics, go back to the very early papers

Specificity & Evidence

7 / 20

The conversation is overwhelmingly abstract and philosophical; concrete anchors are limited to anecdotes (mid-80s Toronto, Duke teaching, a researcher insisting on ANOVA) with almost no quantitative data, named studies, or dollar figures to ground the claims.

the whole university, all the statisticians and biostatisticians were very upset at me and thought I was being a complete fool for showing people how to do meta analysis
if the model is correct, the confidence in low coverage is exactly 5%

Conversational Craft

8 / 20

The host demonstrates genuine intellectual range - invoking Hume, KL divergence pedagogy, and cross-referencing prior episodes - but consistently lets the conversation drift into long mutual monologues about academia and publishing without pressing the guest for harder evidence or productive disagreement.

And on the issue of the future orientation. So for example, we have, um. Many people are familiar with the problem of induction that David Hume talked about
Do you think science has become too statistics focused?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A71%
  • Speaker B29%

Most-used words

statistics58world46science35learn33model26learning24first21scientific21data21different18reasoning17models16back15expectations15representation15bayesian15

Episode notes

Keith O'Rourke | The Logic of Statistics Dr. Keith O'Rourke talks about the logical reasoning behind statistical modeling. Topics include mathematical vs scientific reasoning, whether science has become too stats focused, and vice versa. Watch it on... Youtube: Podbean: Topic List: 0:00 - The logic of statistics 0:30 - What is scientific statistics? 5:15 - The logic of statistics and CS Pierce 9:15 - Role of representation in statistics: explicit vs implicit 14:13 - Diagrammatic Reasoning 18:45 - Why is modeling counterfactual? 19:33 - How can statisticians become better scientists? 28:40 - Science is hard 31:24 - Computational approaches to learning 42:00 - Learning through metaphor 46:28 - Diagrammatic representations vs math 48:40 - Is science too statistics-focussed? 59:35 - Is statistics sufficiently science-focussed? 1:08:40 - Scientific Debate #statistics #datascience #science

Full transcript

1h 13m

Transcribed and scored by The B2B Podcast Index.

Speaker A: What I call logic and statistics, I know I didn't define it very well, can be done without a lot of, um, without a lot of mathematics. Um, and it can be done for maybe simple to almost intermediate problems. So you're not stuck with just having to talk about normal distributions and very convenient assumptions.

Speaker B: Foreign. Welcome back. Today's guest is Keith o'.

Speaker A: Rourke.

Speaker B: He is part of uh, o' Rourke Consulting and also Health Canada. And this conversation simply reflects his views and not those of his employer. And what we're going to be talking about today is the logic of statistics in a very deep field. I'm very excited about this, uh, in some of our conversations. Keith is the first person who's brought up ah, CS Pearce in casual conversation before, which is very exciting. Um, and um, there are a lot of interesting things up ahead. So Keith, welcome to the show.

Speaker A: Well, thank you. I'm very pleased to be here. Thanks for inviting me.

Speaker B: And so I guess maybe we should just start off with the question, what is scientific statistics to you?

Speaker A: Uh, to me it's uh, the purposeful use of statistics to learn about the world in the future tense. So we want to kind of form uh, expectations or a schema of what will happen in the world when we do X or Y or Z. And we hope very much that when we do that, ah, we're not going to be frustrated, like we're not going to fall off a cliff or be eaten by an animal. And uh, one of the things I've been actually reading up on while listening to a bit the Carl Frieston and I don't know if you know about his um, free energy principle, which really is actually very statistical. Um, and he kind of makes the point that any organism to survive has to isolate themselves from the environment, but also represent the environment somehow by receiving signals and then use that representation from the signals to decide how to act on the environment so that they get to eat rather than being eaten. And uh, it's kind of maybe still fairly controversial, but it is quite interesting. Um, and um, with the point, one of the more interesting points is that uh, if organisms survive they've actually somehow figured out how to implement uh, scientific statistics, in other words, dealing with the world and the uncertainties out there in a way to survive, um, which is kind of interesting. Um, so I think really what it is, it's purposely using statistics to uh, create expectations of the world that won't be frustrated when you act. And you know, you could think about descriptive statistics as sort of, you know, what happened last year, which is kind of different. It's why I use that term. It's in the dead past. It's over and done with. Um, so it's this future orientation that's really important.

Speaker B: Yep. And on the issue of the future orientation. So for example, we have, um. Many people are familiar with the problem of induction that David Hume talked about, where effectively we have these observations and as you call it, the dead past. Um, but what we care about is what's around the bend. Um, and we have no logical guarantee that what it was like in the past will continue to be the same way in the future. Um, and so there's that disconnect. Is this part part of that issue that you're talking about? Are there other aspects as well?

Speaker A: No, I mean, that's one of the things that has to be dealt with. And you mentioned, uh, Charles Sanders first. Uh, and he kind of pretty much set that aside saying, well, it really doesn't matter. You have to inquire and you have to act and you have to continue to inquire. And so, um, you know, the past will not be like the future. It'll change, hopefully not too quickly, but you have to keep inquiring. So you just, um. And there's really no alternative. You know, you either have to try and uh, form expectations of this world that you won't be frustrated with. There's no alternative. So if. If the world changes too quickly in a certain area, you might go extinct. Can't do. There's no alternative. So you just have to try.

Speaker B: So, uh, then I guess what would be the logic of uh, statistics? I guess what we want to talk about is the logic of statistics. So that's scientific statistics. What is the logic of it?

Speaker A: Right. And I think pretty much everything I think of a lot of. What I think of, uh, draws from, uh, first, um. And to him, M. Logic was not a narrow field. A lot of people think of logic as sort of formal deductive reasoning. To him it was sort of, uh, any way of um, producing good inquiry, you know. And again, mostly learning. It's not just, you know, more generally, not just learning about the world, but also learning about abstractions that we make. Um, and it's really just the process of learning. And if you're learning about an abstraction, you can do it deductively. So there you want to, as you said, represent and re. Represent in ways that won't take you astray. So your re. Representation has to be, uh, completely valid. Uh, but we reason all the time and we even have the thing he had the biggest problem with, and he never really resolved in his career or lifetime was, um, what he called abduction, which is coming up with hypotheses. Hypotheses. And he liked to think it was a logic, a very weak form of logic. Uh, and he tried to justify it, but in, in the end he never quite succeeded.

Speaker B: Was he the originator of the term abduction?

Speaker A: He might have been, but he also used, he always uses a number of words. So he. He called it abduction, retro direction, introduction, hypothesis, guessing. Um, I mean, one of the things about, uh, Peirce which is worthwhile maybe knowing is that he never like, um. I forget the name of the philosopher who said that he started everything and finished nothing. So what Peirce did is he kept on reconsidering things over and over again. And every time he went back he would bring up further considerations, but he'd never really get it to closure. And I think he really felt that trying to get to closure was a mistake. That really all you can do in science, um, is decide to take pauses until you can feel the opportunities good again. Or what he called the economy. Research. The economy of research is worthwhile to try and take it further. So anyone trying to read Peirce, he's very frustrating because, uh, you know, he doesn't get the closure. If he ever does get the closure, if you come back and read him again, he'll go somewhere else. So, uh, but he's worthwhile. Actually, it's Deborah Mayo who put it nicely once. You don't read Peirce for answers. You read Pearse for inspiration to get sort of, as I would put it, find out about the things that he considered and how he took those considerations through and then that'll help you do better science or better philosophy.

Speaker B: Um, and just for those who aren't as familiar, Peirce, uh, I think he was the first person to propose that we, uh, could use, uh, for example, electrical circuits to perform logical operations and things like that.

Speaker A: He did that. He had a graduate. It's a cute story. He had a graduate student. They used to have logic pianos, right? Certain keyboards would work out certain syllogisms, but you had to have a really good carpenter to build it. And one of his graduate students had problems getting a good carpenter and one quit. So Perth said, well, why don't you just use electronic, um, switching circuits? So, um, that was kind of the instance.

Speaker B: Yeah. So, yeah, ah, really interesting ideas. Um, on the issue of. We talked about reasoning and abstraction. Um, one of the things I want to talk to you about was the role of representation in statistics and eventually get around to this idea of diagrammatic reasoning in statistics. Um, so, ah, what is the role of representation, statistics, and how is that essentially bringing together the mathematical formalism, the real world, how does that all get funneled together?

Speaker A: Well, the mathematical abstractions, they're abstract. We make them, uh, and we make them as we make them, and they are what they are. Uh, and you can learn about what they imply. You can learn, uh, about everything that's contained in that representation with that abstraction you've made. But if you're going to try and learn about the world, you have to realize that you're using a model as a representation. It's supposed to kind of connect somehow with the world. And ah, so then, you know, when you've kind of worked through the implications of the abstraction, you then think they're kind of, they're going to be similar or analogous to how the world works. But somehow the mathematics has to connect with the world if you're going to be learning about the world rather than just about the abstraction. And most of what we do in statistics, uh, everything that's done in mathematics is just learning about the abstraction. And in statistics, a lot of the work we do, we're just learning about those abstractions. And then somehow we have to transport that understanding to expectations about, uh, what will repeatedly happen in the world. Right. If you have expectations of what's going to happen, that's a sense of what will repeatedly occur. Right. We don't expect the same thing to happen every time. Like you go back sort of back in the 1800s when they started taking photographs, uh, of people. They would do things like, um, take photographs of all the, uh, the female actors in movies and overexposed them on the same picture. Uh, and what they got was kind of an average picture of uh, a, uh, female actor at that time. That's kind of the average, um, female actor. That makes sense.

Speaker B: You know, you go, yeah, so far.

Speaker A: Yep, so far it's kind of like averaging. Yes. And now today I think we should be thinking about the distribution, right? And I think, you know, this is kind of, we all have these distributions about expectations of things like if we're watching a new movie and there's a new female act, leading actor and you haven't seen them before, you'll have a set of expectations, the distribution of things that they'll look like. Okay. And so that's kind of very statistical, right? It's a distribution and it provides you a set of expectations for how to confront the world Right. And if, if that female actor looks very different, your distribution of, uh, past female actors, you're going to be very surprised. And, um, I don't know if that makes sense or.

Speaker B: Yeah, I guess so. I guess connecting it back to the, uh, representation. So I guess the, um. I guess. Yeah, go ahead.

Speaker A: So usually, uh, people talk about assuming models when they're doing statistics, Right. And I think sometimes they're thinking about. But you have to assume models is just, uh, a cost to getting quantitative results. You know, something. It's like paying taxes, something you have to do. But if you think about it, what you're really doing is you're representing the world somehow. If you say, um, uh, you have a model with a parameter space, um, well, you're kind of saying, well, those outcomes can occur in this world. Right. That's why I'm assuming this model. And so really you should be thinking about it as a representation. The representation, something that stands for something else in some aspect, for some purpose. Right. So if we're doing statistics, we're trying to learn about the world, okay, we assume this model, but we should realize that we're taking that model as a representation of the world for some aspect. Right. So, um. But people even don't like the word sometimes. Using representation as a synonym for models

Speaker B: is that where, for example, like diagram, like diagrammatic learning comes in?

Speaker A: Well, I, I think we can separate. I mean, diagrammatic, uh, learning is just a way of learning about mathematical abstraction. But if you think you just, you have this abstract model, it's mathematics, got nothing to do with this world. But if you're using it, uh, to learn about this world, well, then you're using it as a representation. It's sort of standing in place of the actual world. Right. It's like this idea of a Markov blanket where you have to separate yourself from the world. It's really dangerous to bring things into your brain. Right? Right. So somehow in your brain you have to represent the world. Right. And in statistics, we're doing that with models, in particular probability models. Um, and so, and a lot of it comes from, like. I started out a long time before I took statistics. I studied something called semiotics, which they call the theory of science, or better today, theory of representations. It's how something comes to represent something else for some purpose. And so it was really easy for me to see. Well, if you're making assumptions and statistics, well, then you're kind of taking something as a representation because you're assuming those models because you want to learn about the world, right? And so if you're doing that, you're essentially using as representations for the world as stand ins for the world. So one of the ways that, you know, we're talking about this, and I think Andrew Gelman has as well, and we had this difficulty whether we call them fake worlds or whatever, but you can think of this, you know, you can think of these, um, probability models, okay, as fake worlds. Now, because it's an abstraction, you can learn everything you want about that world, right? You can learn all the implications, okay? Whereas in the world when you observe some data, you can't observe it again, right? You can't put your foot in the same river twice, right?

Speaker B: Yeah.

Speaker A: So but if you use the probability model and say that data could have been generated by this probability model, right? In an idealized sense, okay? Well, now you can learn about what would repeatedly happen in that fake world if it was true. And then there's a hope that that fake world is close to the actual world that you have to deal with. And so, you know, what would repeatedly happen in that world. And so then you can use, you know, kind of transport that to what would repeatedly happen in the world. I mean, one of the things that allow you to do is it sort of allows you to work out, um, how compatible the observation is with that probability model, right? If it's in the tails of that probability model, something that you'd seldom observe if that model was true, okay, well then it's not very compatible. And there's various people that have worked this up, like, uh, Sandra Greenland, and she worked it out into sort of compatibility assessments and compatibility intervals rather than confidence intervals. And the key here, the point I'm trying to make is to get that you have to assume that a fake rule is true. That's abstract. Now you can learn everything about that. And now you have to relate that to the observation you have. So the observation you have, you can see where it's in the, in if it's in the tails of that fake world. Well, it's not, uh, that compatible with that fake world, right? And if you change each time you change, let's say for instance, the mean, well, you're changing the fake world, right? And so you can work through basic statistics like that.

Speaker B: What about the, uh, issue of counterfactuals that you brought up that I thought was very interesting and it seems like it's spot on. This is about as good a time to bring it up as any or effectively that the probability distribution is something like a counterfactual yeah, it allows you to fact.

Speaker A: Because what's counterfactual is what you would observe next time, which you can't do.

Speaker B: Yeah.

Speaker A: Uh, because the world's going to change even second from now or whatever. Uh, so it allows you to get at that counterfactual. Right. What would repeat, what would be repeatedly. What would be repeatedly observed in the world. Okay. You go to a probability model abstract. And you can work it out there and then you kind of have to like transport it to your observations.

Speaker B: What do you think? Um, what do you think sort of the best takeaway for you know, these early stage young statisticians and data scientists would be from these sort of ideas? Um, because you know, I think one of the challenges is that when people are just getting started out in the field, they have a lot of the, essentially the mathematical, the mechanics inflicted upon them, um, and not so much the connection. So effectively they, they become, it's like. Well, one thing that we can always test people on is we can train them in the abstractions, we can train them in the math. It's very easy to assess the math. What's not easy to assess is if they're a competent scientist. Um, which is especially difficult because, you know, becoming a competent scientist takes a long time. Whereas you can train someone up in the math in about a year. Um, you know, you can teach them to code within about six months and then, you know, get them pretty good at coding. So effectively they can go through the motions of being a statistician or being, going through the emotions of being a data scientist. Um, but then there's all that other stuff out there. What would be a good roadmap? Not a roadmap, just you know, what you're talking about. What would be a good takeaway message for young people?

Speaker A: I mean actually thinking a bit about that this morning, probably, you know, my sense would be, uh, ideally they would do this stuff in a pre statistics course. You know, anyone's uh, going to be a professional statistician. There's a lot of technical math you have to learn. It's different in different areas depending on where you go in. Um, but you really, if you're certainly, if you're going to be doing research and statistics or really trying to be, uh, you know, the best or the best, trying to be a best statistician you can be. There's a lot of math you have to do at some time. I think a lot of these ideas can be done before that. Um, and you know, it's. We're Getting ahead of ourselves. But I think I sent you some material where I use um, simulation and then I point out that you know, Bayes can be done with what's called direct or two stage or ancestral simulation. That doesn't take a lot of technical skills. And then you can bring in. Sorry.

Speaker B: Oh, go, go, go on. I was about to. Interesting. Bitcoin please.

Speaker A: Yeah. And then you can bring in uh, like Bayesian workflow, like uh, Michael Bettencourt and other Andrew Gelman's colleagues work with where you kind of interrogate how well your Bayesian analysis worked. Right, right. Because if you know with a Bayesian workflow you're going to have to do the same ancestral simulations anyways. Even if you can do MCMC sampling, you're going to have to do this uh, two stage simulation to assess whether that worked very well. Because you're going to have to assess first does the prior that I've used, that is, does it, you know, does it imply marginal priors that make sense to me. So you have to stimulate from the prior to sort of check the prior and then you have to stimulate from the prior and then you have to simulate the data given the parameter you got prior to see if your Bayesian model generates um, observations, potential observations with fake observations that are like what you see in the world. So you have to do those steps. So you know this, this whole, what I call logic and statistics, I know I didn't define it very well, can be done without a lot of um, without a lot of mathematics, um, and it can be done for maybe simple to almost intermediate problem. So you, you know, you're not stuck with just having to talk about normal distributions at very convenient assumptions. And one of the ways that the Bayesian workflow uh, works is you say, okay, I'm gonna uh, my data generating model, let's say my prior is going to be normal and my data generating model is going to be normal. But what you do in the Bayesian workflow is you change that. So you make the prior something else and you make the data generating model something else. And you say well this could be more like the world. I won't know that. So I'll use my normal, normal Bayesian model to try and learn about it. How well do I do? And the thing that saves us is the statistics you should usually do. Not too badly. Right. There's some mismatches that will get into a lot of trouble, but you can get away with a lot of miss, you know, not representing the world that exactly. And still kind of, you know, form uh, expectations that reasonably are well met. Um, so that we've kind of got a little bit ahead of that. Now one of the problems when you do that with people that don't have a lot of statistical knowledge, then they don't appreciate it because it looks too simple from them. Too, too simple for them. Right. I'm simulating from the prior, okay, I got that parameter value simulate the data. I do that over and over again and then I just keep from. That gives you the joint distribution, but then I just keep those of the joint distribution that match my observation and that's the posterior. Because other people have used this to teach novices Bayesian statistics. Um, Richard McElhreath is one and uh, uh, a few other people, they say what happens is the students find it just too like uh, trying to learn skiing on the bunny hill. They want to get, they want to go down the steep slopes right away or they want to learn to do real statistics. MCMC sampling and often I think most introductory Bayesian courses, uh, almost all the time is spent dealing with MCMC sampling and there's no time spent really. Well, how do you choose prior? What's the downside of choosing it badly? We also have to choose the data models. And how would you ever. I mean it's all blind, right? The usual course you usually say about use a default prior. They work okay, um, and then use your regular data turning model. Usually assume, um, run the mcmc, spend a lot of time checking that it's converged and then the posterior is your answer and it's totally blind to this sort of scientific reasoning. So the scientific reasoning is, you know, you've represented the world in a certain model, you've worked out the implications of that model, you've assessed the compatibilities with what you've learned from your, from all the implications of your models with the data in hand and then you decide whether to go with it or not.

Speaker B: Do you think. So the disconnect between that is, uh, between essentially the, I wouldn't exactly call it Bayesian workflow, but just for that, like essentially their statistical workflow. Do you think that there's some disconnect where effectively anyone can learn how to turn the crank on their statistical workflow, uh, but how to actually turn the crank on your scientific uh, workflow like there is no one crank to do it, you know, where effectively the science, the scientific reasoning is the big nebulous answer or the nebulous question. So effectively people can operationalize one aspect of it, they can't operationalize the other. And therefore they choose to sort of really dig in on the one thing that they can't operationalize. Is that a fair assessment or am I missing something?

Speaker A: Um, I'm not really sure because there's a few things that come up here. One is like I did this material with my son when he was a teenager and he looked at me, says dad, I get it, I see what's happening. Two stage sampling is pretty easy to understand. But shouldn't there be a formula that does a better job?

Speaker B: Right.

Speaker A: I think that's part of the disconnect. The sort of expectations that math uh, is always better, um, and you should be doing math um, the like. Science is very hard, it's extremely hard. Um, and I think there is a temptation to just um, use what you learned in graduate school, which is mostly just the mathematics and just take that as providing your answer. Um, I remember something talk I saw a few years ago where someone made this point. I think it's true for most of us in your university training almost no one talks about scientific reasoning. You know they, they kind of mostly just well this is the way we do science or this is the way we write statistics papers. You know they expect this and they expect that. No one really gets into the larger question. Ah, as the speaker on sort of of mentioning said, how do we know what we know? And part of the problem is um, at least in my mind is one of my uh, philosophy colleagues uh, told me once that one of his uh, faculty members described philosophy as asking wild, large childlike questions and arguing over the answers like lawyers. And you know the philosophy has its place but most people don't have the uh, time or aptitude to go through that. And what they really want is something that's going to help them do better science. And I keep trying to think of a better term than philosophy of science. And you know the only thing I come up with is sort of theory of purposeful inquiry. But it's really, it's, it's all this is how do you learn about the world profitably? So you know, profitability in the sense of being able to act with less frustration. Um, and, and that's not taught anywhere. So uh, and you know it's. So to go back to the original trying to answer your question, if we had a pre calculus, pre uh statistics course that taught sort of scientific inquiry and use these simpler like things like two stage sampling and you can leave the prior off and just make a, about classical statistics. Um, and so it's Just a way of um, really focusing on using uh, probability models to represent how the observations could have been generated and then using simulation to learn about those because that's really easy to do. Uh, and then doing this mismatch where you represent something one way and analyze it, thinking you was represented another way. Uh, and seeing what happens and what can go wrong. Yeah.

Speaker B: The issue of uh, essentially using more computational based approaches and simulation based approaches to help students learn, to me seems like I've been thinking about this for quite some time that there are. It seems like it could be very profitable at least for a certain subset of students. Um, where um, if you essentially have that more. That more like interactive aspect to it. It's computational so it keeps it very simple. Um, essentially computational approaches to teaching statistics as opposed to analytical approaches to teaching statistics. Um, and it isn't just in statistics, there are other fields too where I think it can be very helpful. For example signal processing. I think it could be immensely helpful where effectively a lot of the analytical ways that they teach the topic of signal processing tends um, to lose students unnecessarily on certain details. Um, where if you taught it computationally it would lend the topic more to the scientific applications for which it's used. Um, and I think there are plenty of areas like that where if they taught things computationally and I haven't quite connected why that would be other than I think that people just learn by different, have different modes of learning.

Speaker A: Well again, you know, computation is a very general term. Yeah. Um, and um. I think um. One of the things I worry is. But often it could be seen just like a black box. And um, one of the things I try and do, I, I can go through this very quickly. I use digits of PI as a uh, basis to generate zero uniform random numbers. And then I take pairs of them and use rejection sampling to get other distributions. Just so there's. There's a black. Not, you know, there's no black box there.

Speaker B: Yeah.

Speaker A: You know, that's a very inefficient way to simulate from let's say a normal distribution. But it's very easy to understand and then you can use the more efficient ones. So I think you do have to make some effort um, to make it clear how simulation works. Right. So that it doesn't just become a black box. Now one of the problems or one of the challenges going to be uh, if you're going to use probability models to represent more than trivial problems, you know, that starts to become very challenging to like represent. Um, you Know, because you're going to have to put together a bunch of probability models and that's a skill in itself. Um, so. But a lot of them might be driven just by, you know, I remember one year I, I taught at Duke University, uh, of course in introductory statistics. And you have all the rest of the university or different faculty members, they have certain expectations of what their student is going to learn in your course. And so you really don't have a lot of freedom to design your own course in most universities. You also don't have the time. Takes a lot of time. It's far quicker to just use the same old worn out pedagogy that everyone else has used in the textbook with the, the answer keys and spend as little time because unfortunately academics have a lot of pressure. And so, uh, if you're not doing certain things and getting publications and other stuff, and so you really can't afford to put a lot of time into kind of, uh, exploring, um, how to teach statistics computation.

Speaker B: Yeah, it's why I think actually I won't open up that can of worms. Uh, but just as a quick, as a, as a quick rewind, you know. Uh, one thing that I think is worth pointing out though, is that, you know, the analytical approaches to teaching statistics, the moment somebody zones out the math behind it, that becomes a black box as well. Um, and so effectively, I think it's, effectively, you can choose your black box. And to me it seems like the computational approaches are far less black boxy, the steps are simpler. You know, there's no thing where someone just like skips a step and says, here, QED or whatever. Um, go on.

Speaker A: Yeah, no, I mean they're less black boxy, but it doesn't mean that certain people aren't going to take them as black box.

Speaker B: Yeah.

Speaker A: So I'm saying you just, you know, you pull that apart. Um, and, um, I think maybe the other thing would help is to use it in this misspecified model way, you know. You know, uh, simulate, um, data one way and simulated a different way to try and learn about. So there's a mismatch. Um, there's something, I don't know how familiar Bayesian statistics are about this, but there's someone who, uh, was investigating confidence interval coverage of Bayesian, uh, models and they did it, ah, using the same model. They use it using the correct model. So in the Bayesian model, if the model is correct, the confidence in low coverage is exactly 5%. Right, right. Um, but not everyone knows that. Like I've, you know, Various people who, you know, they're doing Bayesian statistics. And I point that out and they go I didn't know that. Um, but that's a very. What you really want to do is misspecify the model so that it's not. And it's 5% on average.

Speaker B: Right.

Speaker A: If you fix a parameter it could be 10% or 1%. Um, but it's, but I think at the bottom it's just science is very difficult and then also learning and teaching is very difficult as well. And it's really hard to know um, when you've given a lecture or something exactly what people are taking away from it. Mhm.

Speaker B: I think there's also a big problem where you can't really begrudge people for becoming fairly mercenary very quickly or they just become extremely practical on the like do I need to know this? And I don't learn anything beyond that because you know like essentially they get funneled through decades of schooling by which they're graded on. They literally have this um, it's a really cheap cost function by which I mean it's like a poorly designed cost function or it's not a, um, you're not doing ah, like multivariate optimization or ah, multi um, ah, darn the words coming out of my head. But essentially you not doing like ah, um, you're not doing a multiple cost function, you're optimizing on one cost function. It's basically like did you learn enough to pass a given test? And there's no reward for knowing things about that. And not to mention the test is being designed to untestable things which is the other, which is the, I think the more fundamental bit. But effectively we basically spend all this time convincing large number of people to be totally mercenary and narrow focused on what the value is and then once they come out into the light of day it's like oh, now all these other things matter kind um, of sort of. Because even then you don't particularly get rewarded on them. Um, but it does seem like um, essentially they get trained, they get over trained if you will. They get trained on this very narrow set of problems. Um, because on, by virtue of need to grade and have this sort of academic system they get funneled through and it's not surprising me that essentially science and creativity. And I think that's another thing that I think science is such a fundamentally creative thing, um, that the same aspect of education that kills creativity in so many people is the thing that kills their scientific capabilities as well. Like I think those are Twins, and they get. They both get poisoned.

Speaker A: And I think you're right. My daughter's okay with this. When she went to undergraduate university, I would talk to her and say, don't focus so much on getting high marks. Focus on learning important things. Um, of course, um, you didn't believe me. It's normal for kids to do that. And it was a few years after she got out of the university, and she told me how you were right. She'd been focusing on learning important things. And for some reason, myself, I did that in university. Um, I know it used to startle some professors when they say, oh, I'll let you rework on this assignment so you can get an A. I think I had a B or B plus. And I said, why in the world would I want to do that? You know? And when I was in undergraduate university, uh, actually, most of the time university, I was always going to different seminars in different departments and, you know, reading other stuff and just kind of following my interests. Um, but, uh, most people, you're right, Glenn, they're trained. And like, I experienced that, you know, when I was teaching at Duke University, is how do I get whatever grade I want to get with different targets with the least amount of effort? And they were very upset if you made. If you made them think through things or work through things or try to understand something other than just kind of guide them, um, with a quick route to get their B plus, and then they go on to other things.

Speaker B: Um, so, uh, obviously we've talked quite a bit about the challenges of learning statistics, and I think one of the things that you and I both agree on is the value of, for example, metaphor, or I like to use analogies when learning about concepts, just any concepts in general. Um, why is it that you think that metaphors are particularly useful for learning?

Speaker A: Um, well, I like them. I mean, the basic idea is you, uh, transport an understanding in a familiar area to a less familiar area. Um, and, uh, we could go into the semiotics, but that would take us the rest of the hour. But they're very interesting objects. Uh, and, uh, there's a few that I think have worked well for me. Um, one is actually kind of a joke. It's a Picasso joke. And it's, um. Picasso painted a picture of this husband's, uh, wife. And the husband met Picasso at a party and said, picasso, that painting of my wife does not look like her at all. And he goes, oh, really? What does she look like? So he takes a picture out of his wallet and shows it Picasso here, she looks like this castle goes my. She's awfully tiny.

Speaker B: Mhm.

Speaker A: So that's kind of a metaphor because people, when I talk about, you know, probability models as representing how the data were generated, you know, the first impressions was they're not really like that, or it's not exactly like that. Well, representations are kind of, they're like this little painting. They don't have to get all aspects right, they just have to get the important ones. Uh, the other one I like is using uh, shadows as a metaphor for statistics. And that's so you see a shadow but you can't look at what's casting it. And what we're doing in statistics is we're trying to learn about what's casting it. We don't care about the shadow per se. So that's like uh, the observations in our study there, the dead past. We really just want to use those to learn about what will repeatedly happen in the future. Um, the other one that the ah, chemists like, because they actually do this, is they spike uh, known amounts of chemicals and test tubes and then they take machine readings and then they sort of calibrate the machine readings, um, sort of, they're really working out the distribution and machine readings for the known amount they put into the test tubes. And that's kind of what we're doing in statistics. But we have to use probability models because we can't spike uh, certain things like cancer rates or return on investment. So we have to make these uh, probability models to do that.

Speaker B: Do you think that, uh, one of the reasons, getting back to, you know, the, the conversations about um, you know, diagrammatic reasoning and representation is that are these metaphors helpful? Because effectively it's easier sometimes to manipulate the metaphor in our mind to make progress.

Speaker A: Um, no, I think metaphors work more like schemas. They create sort of expectations of a certain area. Uh, and that's why that was something I mentioned about the semiotics is that uh, or even with the Friston material at the beginning, I think our minds form expectations of stuff. So a metaphor is a way of implanting expectations. Right? So you have these great expectations about shadows and now you want to sort of set them up for statistics. Uh, and so it gives you a set of things to expect in statistics that you wouldn't have expected before. The diagrammatic reasoning, as far as, um, I can tell that's just a medium of mathematics, right? Because if you have an abstract quantity, those are mathematical quantities, you won't have any other kind. Uh, you Want to discern ah, the forms of relationship in that. However you do that, if it's just however you do that, as long as it's a reliable way of doing it, that's mathematics. So whether you're doing it with symbols or with a diagram, um, you know, you can, you can do diagrammatic proofs. People used to think those weren't very good, but you know, they've since argued that can be as accurate as deductions, you know, if you do them correctly enough.

Speaker B: Um, can you talk about that a little bit more? Why is why however, for example these. Well, why do people think that this wasn't as good as I guess, mathematical deduction?

Speaker A: Well, because I think people did poor diagram. You know, they do diagrams loosely and would draw incorrect conclusions. So you can do that with any form. You can do that with any form of mathematics.

Speaker B: Mhm. So basically does mathematics in many ways traditional mathematics, it's very good at forcing you to be very rigorous and precise. And when people then jump to a more diagrammatic approach that they might lose the precision or what.

Speaker A: There's arguments that you can make the diagrammatical approach as rigorous, but that's not really its feature. Its feature is it is just easier for most people to understand. Again, like anything else from uh, what I'm working on comes from Peirce who defined mathematics as performing experiments on uh, diagrams or symbols instead of chemicals. And so it's experimental reasoning, uh, performed on the abstraction and the diagram just makes it clearer that that's what you're doing. So um, you could sort of draw any arbitrary triangle and prove it has 180 degrees by sort of just manipulating lines and checking angles and you'll do it. And so that's kind of a diagrammatical approach to proving that all triangles have 180 degrees.

Speaker B: If you don't mind me just hopping back onto the uh, the science for like ah, the, the scientific aspects. Um, do you think that given some of the challenges and the turn the crank aspects that practicing statistics has sort of picked up as its burden? Do you think science has become too statistics focused?

Speaker A: Yeah, I, I think if you let me first answer it a little bit differently. I think it's too single study. You know, the idea uh, that you can learn anything from a single study is to be a little bit bizarre. Occasionally we may be in a situation where we can only have one study, uh, but usually there's going to be more, usually half a dozen or a dozen. And so I think it's a big mistake in statistics that for Some reason everyone thought you could just focus on one study at a time. It's a big problem in academia where authors are expected to come to conclusions based on their one study. Well, this is one of maybe a dozen studies. No one should be trying to get conclusions from one. And, um, one. When I first started in statistics back in the mid-80s, I was at, uh, the University of Toronto, and I started working with some clinical researchers and I worked out for them how to do meta analysis. I had just learned, uh, generalized linear modeling. And so the studies we're working on had binary outcomes, so that's logistic regression. And you put an intercept in for the study. So each study allows a different control rate. Patients can be sicker in different centers. And then you look for a common treatment effect and then you check if there's an interaction between treatment effect and study is sort of the heterogeneity question. It was very straightforward to work out and then we started doing it and then the whole university, all the statisticians and biostatisticians were very upset at me and thought I was being a complete fool for showing people how to do meta analysis because they should know that you're not supposed to try and do that. Um, and there was this general pushback, uh, for years and years about looking at more than one study at a time. And you shouldn't. And, you know, all you should really do for a single, uh, study is, you know, do your best to represent what happened and what you observed and then let someone else come along when there's a half a dozen or so studies to actually try and draw conclusions from it. And the other one is more to your, um, pulling the crank. There's, uh, too much use of convention, right? Um, we need defaults, but we want to avoid just conventions. I mean, um, Andrew Gilman and I wrote a little commentary, maybe it's 10 years ago or so, about how statisticians choose what statistical method to use and the biggest predictor is where they went to graduate school. Well, that's what I learned to do in graduate school. So that's how I'm going to do your analysis. Um, and, uh, I remember I was talking to a young, um, researcher a few years ago and they had done this really nice experiment and the results were actually kind of clear. And I was talking to them about how they could display it in the graph and make it very obvious what happened in their experiment. And they kept on, no, no, no. How do I put it into analysis of variants? How do we put an, anova table? How do I get an ANOVA table? And I kept trying to tell her, do this first. I'll put it into an ANOVA table to control people, but you should do this first so we understand what's going on. And m. They never got back to me. I figure they probably bumped into a statistician who could show them how to make it into an analysis of variance and that's what they did. Um, but you know, it's, you know, people like to be sophisticated and they also use all these conventions in uh, academia. Like if you submit a paper in your area and they're expecting analysis of variants, you're going to maximize your chance of being published by putting it in like that. So uh, I think it's um, it's maybe default thinking and too much convention. When I was actually first uh, taught to do applied statistics, it was actually by David Andrews at the University of Toronto years ago. And he always taught us to uh, think the problem through by first principles. Don't look in the methods book, don't look at how other people analyze it. Just think about the problem and what the experimenter is trying to learn and what they did and what they observed and try and you know, make an analysis from first principles. Uh, and then of course before you finish, do check the literature to see what other people are doing. But they do the other first and, but that's hard. It's hard. It's time consuming. Sorry.

Speaker B: I would say that's really interesting because um, I've been writing up some lectures on KL divergence. And the way that I wanted to teach it was not by jumping immediately to the answer, but essentially saying, okay, if we just started with this basic problem, how would you off the top of your head, propose to resolve it? And we can try it this way. And then essentially we work our way and see what the implications of that metric are. And then I propose an alternative metric and we work it and we get out to uh, some uh, von Mises criterion, uh, or von Mises metrics, um, um, and then uh, doing it another way and we get to some like Mulgraf Smirnoff test and essentially we try it different ways and see where each of those tests go. And what I attempt to demonstrate to anyone who would listen on this would be that um, eventually after we cure the shortcomings of these other methods, um, or essentially correct them, we get pretty darn close to uh, Kullback, uh, Leibler divergence. And it's like, okay, actually if we just sat around and tested and played around with enough we could actually produce something pretty darn close to what these geniuses in our field came up with, um, if you just kept working at it. And um, so I thought that was interesting. And my hope was that by spending the first, say 10 minutes on a lecture doing this, um, I was first 10 minutes, a three part lecture. The first part just coming up with this is that then it's not some magical thing handed down to them. They're actually very comfortable with the concept and from that they can have a more profitable engagement and they can use it more dynamically. It's like a Swiss army knife where you've put it together yourself.

Speaker A: Um, I think it'll be a great benefit to them as long as they'll go along with it. That's, that's a bit of a challenge, but I think it's certainly uh, a great way to teach uh, an understanding of statistics as opposed to just this is how it's done, this is the math, get used to it. You know, that's the, that's the other expression from one of the mathematicians used once you don't understand math, you get used to it. Well if you're going to use math to do science, you better understand it because you have to. Um, but uh, it's uh, the other thing actually once years ago Don Rubin told me is that whenever you're learning a new area in statistics, go back to the very early papers.

Speaker B: Right.

Speaker A: Because in the very early papers they were struggling with how to do it and they hadn't slicked up the mass so you can't understand what's going on. It's very similar to your planned teaching uh, lesson because the early literature would be kind of doing these uh, sort of haphazard walks towards something really spiffy, but they're not there yet. So you'll understand the spiffy method because it was m. Spiffy method resolved all these earlier efforts. Mhm.

Speaker B: Yeah. It's another thing also, you know, you've heard the Ruta Jung episode, uh, which will have come out by the time people uh, hear this interview. And you know how he talks about how at the beginning when people are really breaking like really forging a path, they're really thinking about the reasoning behind things so that they're the ones doing a lot of uh, thought and then the, and it's not actually formalized in mathematics yet, it's just that intuition based learning. So effectively they have intuition, creativity, they're charging ahead and it's only once They've actually found something useful that someone will come on. It might even be them at that point. But you know, then they come along and clean up and figure out the math. Um, but the problem in following these people's footsteps is you get so swept up on having to learn the math and all the technical components to truly understand it that you've lost sight of the reasoning behind what they did in the first place.

Speaker A: The best example I had was years ago. I was at graduate school. We're studying for, I think it's a midterm in experimental design course. Uh, and we're talking about interactions. And so I said well I could make a two by two table with an interaction in it. And they go what do you mean? Okay, so the treatment effect is different in A plus versus B plus. These are the means. Right? So the means are different. So that's an interaction. And they go um, what are you talking about means? This is analysis of variance. So they didn't understand that the analysis of variance. They were getting good marks in the course, but they didn't understand analysis of variance was about uh, looking for differences in means.

Speaker B: They totally missed that it says variants in the name.

Speaker A: It's not. Yeah, and I like they've been doing well in the course. They could do all the formulas.

Speaker B: I guess this uh, is teeing up to my next question. Um, well I guess I think we've teed it up pretty well. Um, and we've probably answered it by the time I asked. But um, uh, is statistics sufficiently science focused? So previously is science too statistics focused and now it's. Is statistics sufficiently science focused? No, no, no. Oh yeah.

Speaker A: And I think it's, I think it's improving.

Speaker B: Mhm.

Speaker A: Right. I mean when I took some of my stats courses, I GUESS it's about 30 years ago it was like proof, lemma, lemma, proof, lemma. Whereas now I think you know, most of them have a little bit of science. I mean obviously not enough. Um, and people are doing more to sort of uh, work with actual data, uh, and do sort of more lives like examples. But again there's no really uh, there's no real discussion of what is scientific method or scientific reasoning. I think um, you know, if we could survey or actually better audit statisticians working statistics jump in, drop into their office and ask some pointed questions they have to answer. You know, people might think things like science is objective or science is a collection of known facts, really naive views, um, and you hear a lot of them in the uh, frequentist phase, uh Battles about, you know, uh, not making assumptions or, um, not so, um, not bringing in judgment. You know, you shouldn't, shouldn't be bringing in expertise. You should let the data, the data speak for itself. And I like the line that Sandra Greenland use for that. Data speaking for themselves. If you think data are speaking from yourself or themselves, you need help from a psychiatrist, not a statistician. Yeah, but it's not their fault. I mean, there's not any, uh, there's not really any space to learn more nuanced senses about, you know, scientific methods and scientific reasoning. Um, and most philosophy of science courses, if you went into them, you're not going to probably get much out of them. They have a different, they're there for a different reason.

Speaker B: Yeah. One of the biggest, uh, things that I would like to see go by the wayside is this idea that, uh, science is a, essentially a collection of known facts. I think that is probably one of the most poisonous misrepresentations of science. And the thing is, it's not even a fair one to have because I don't think you can read a popular science book that doesn't dispel that in some way. Ah, ah, ah, Carl Sagan, um, Demon Haunted World. I was, uh, reading that a few months ago and you know, he basically dispels it every chapter. There's nothing that Richard Dawkins says, for example, that he doesn't knock that one out. In a way. Um, a lot of just popular science, you know, they focus on the methods and things like that, um, and the discovery process. Um, but it gets said so much that I think that anyone who has done any amount of reading and even very, very tractable publicly available popular science books should have heard this. And I just don't understand why people haven't.

Speaker A: Um, well, especially now with the pandemic, there's a lot of conversation on the radio if they're listening to the radio, about science evolving and that if it's changing, that's a good sign because they're learning. Uh, but the other one is Brian Ripley once said, I'll quote him, so it won't be insulting people myself, statisticians don't read. Most of them don't. Um, uh, and that's the other thing you'll see a lot when people reference your papers. Um, often they haven't read them. Maybe they've scanned the abstract. They probably ended up there because when the reviewer said you should reference so

Speaker B: and so, or better yet, if you were the reviewer, say, oh, you really ought to Reference Glenn on this one. Uh,

Speaker A: that's hard. I mean, I've been in position a couple of times, you know, talked to the editor. I said, look, you know, uh, I don't know of anyone else to reference other than my own work that will actually deal with this specific issue. Uh, most editors are okay, but you really have to be careful. Like you have to have a good reason to reference yourself.

Speaker B: Yeah, uh, I know I get sort of pigeonholed. Ah, where there's a. Within ieee, I get like not a small number of uh, papers very specifically related to essentially the intersection of. We'll just put it as like biological time series, patient vial sign monitoring, uh, Bayesian nonveriff parametrics work where I get a fair number of offers and they come so close. Um, unfortunately, um, I always get a little bit tickled when someone does reference me and I happen to be one of the uh, reviewers on it. Um, just as a disclaimer. Of course. I've never actually suggested that anyone does it, but I do find it funny that the idea that someone would uh, use their position as a reviewer to get more citations. Um, but the nice thing is you do actually get a good, uh, this is going in too much of a can of worms, but there you could get a very good dialogue if the, if the, literally, if the review process were more functional than what it currently is, you could have way better dialogues. Um, because effectively when you really know what someone has like the. Nevermind, I was about to open another can of worms, I'll move on. But yeah, it is something interesting where um, there is an interplay between. Once people know to think of you in a certain field, you do end up getting a whole bunch of review requests for things in your, in your niche.

Speaker A: Yeah, I've been. I generally just say no because I'm not in academia anymore, so I don't have that. Um, if, if I review a paper, there's no kind of, uh, no one has a checkbox. In fact, it's almost the opposite. Um, having said that, there's a few. What's really satisfying when you review a paper is when you review one really critically and then the authors come back and they've really improved the work. Um, and there was one editor once in contact and he said, ah, the authors just didn't do what you suggested. Are you still up for reviewing it? And I said yes. And the authors came up with something much better than I suggested. I was really happy, let the editor know that the paper was really good. Um, so there can be Rewards of reviewing. But, um, I generally just don't because there's no advantage to me right now.

Speaker B: Um, I've been talking to somebody who's looking at, uh, I won't talk about the idea too much, but essentially to get a better, um, reward mechanism, to bring essentially, ah, very engaged, functional reviewers back into the scientific review process. Um, that I think would be very helpful. Um, but yeah, no, I'm glad that there's a positive story because very frequently what I experience is you try to give people very good helpful feedback essentially to buff up the scientific validity of an experiment. And the goal of many authors in the face of review is what is the minimal I can do? Or worst case, oh, I'm not going to win anything here. I might as well skip over to another journal, uh, that might not have had the author essentially. So it's like if I try enough journals, eventually one will have three authors who aren't paying any attention and will let this go through without any, uh, edits.

Speaker A: Yeah, academia has a lot of challenges. A quick one that Mike Evans pointed out to me, he reviewed a paper with once and the author referenced something and Mike was interested in the reference. So he read the reference and he said reference, uh, goes back on his review. The reference clearly shows what this author is doing is wrong. So they rejected the paper. A year later, he seems sees the same paper unchanged in another journal and he looks at the references and that reference that Mike used to find out was wrong is missing. They just admitted it and got it published that way. It's a shame. I don't know how they're going to fix academia. It's a huge, uh, challenge.

Speaker B: But, um, well, I guess maybe that leads to, um, the next question, our penultimate question. Um, what is one topic that you'd like to see statisticians debate?

Speaker A: Um, yeah, I think, uh, scientific reasoning and representation in statistics have a prerequisite for doing good applied work. Um, I think that would, it might even be time for something like that. Um, you know, you have to remember, I don't know if it was 15, 20 years ago, applied statistics was a bad word. Most statisticians didn't want to do anything applied. Uh, so it's changing and I think, you know, actually I don't want to get into too much, but sometimes these changes are driven by what feedback we're getting on our work. Um, and the most important feedback is a reality. When it complains, it just frustrates you. And then people learn that, uh, they're not as smart as they thought. They,

Speaker B: and I guess our final question for the day is, um, what is one topic that you would like to see the scientific community debate?

Speaker A: Um, yeah, I probably say the same thing. Um, um, maybe because it's science more generally. This issue of, uh, alternative ways of getting a better sense of good inquirer other than philosophy of science, you know, making it into something other than, you know, not omitting philosophers, but taking it out of the philosophy, uh, of science camp and putting it into mainstream science. I mean the speaker I mentioned who pointed out that no one learns scientific, uh, reasoning anywhere in their program, he wasn't a statistician, he was talking about science areas. And uh, his point was you, you learn how you learn the accepted methods in your science at that point in time. Um, and, and then, you know, unfortunately a lot of people in their careers, the only time they really can put a lot into learning newish things is when they're in graduate school. Um, afterwards, they often for various career pressures outside of academia, inside academia, they really can't invest a lot of time in learning unless they're doing it on their off hours and weekends. And not everyone wants to live that way.

Speaker B: Yeah, well, I guess that's one of the reasons why I wanted to have these type of podcast conversations where we could um, provide a tractable way for people to start getting access to these types of ideas. You know, when we started the Philosophy of Data Science series, um, and starting off, um, you know, first thing we do is obviously have to show Andrew Gelman so that people paid attention. And then subsequent to that, um, you know, this reintroduction of, you know, what is the difference between abductive, deductive, inductive reasoning? Why does it matter for data science? It's not some philosophical thing. This is actually very practical. These are fundamental building bricks for us to perform inquiry. Um, and so, uh, hopefully, uh, people are getting something from that.

Speaker A: Uh, I think it's very good. I think it's a great idea. Um, when I was learning statistics, the only time you get this kind of material is if after a seminar or you know, a symposium or something, the speaker would go to the bar and be sitting around the bar talking about these kinds of ideas. There was a. When I was in Oxford in the stats department, David Cox gave a really nice talk for the graduate students talking about these kinds of issues. And at the end one of the students said, well, you should write for that. And he goes, no, no, no, I would never put that in writing. So it's um, so it's great that these exist.

Speaker B: And after which Brian Ripley then would have said, ah. Oh, Statisticians wouldn't have read it in the first place.

Speaker A: Well, there's. Yeah, I mean, there's always that comeback, but, um, I think he was talking about most statisticians, not all statisticians.

Speaker B: Cool. Well, Keith, again, thanks so much for your time today. I really appreciate going, uh, through these ideas, uh, for the first time. And, um, uh, I look forward to having you on again in the future.

Speaker A: Sounds good. Thanks, Glenn. Have a good day.

More from Data & Science with Glen Wright Colopy

All episodes →
  • Jack Fitzsimons | Evil Models: Hiding Malware in Neural Networks
  • Scott Cunningham | Causal Inference (The Mixtape)
  • Eric Daza | Important Ideas in Causal Inference
  • Wenting Cheng & Weidong Zhang | Advances in Biotech/Biopharma
  • Ruda Zhang | Gaussian Process Subspace Regression
Explore the best B2B AI & Data podcasts →
All Data & Science with Glen Wright Colopy episodes →