Five reasons why statistical expertise still matters: The human side of predictive modeling
Automation can build models in minutes. Statistically grounded tools help you understand when to trust the results, question the assumptions, and make confident decisions.
Chris Gotwalt, Ryan Lekivetz, and Russ Wolfinger
July 28, 2026
7 min. read
Each of us must build and maintain trust in any tool we utilize and accept full accountability for results we produce. — Russ Wolfinger
Predictive modeling plays an important role in how organizations make decisions and uncover insights from data. The rise of AI tools has added new dimensions to that discourse, raising substantive questions about automation, skill, and where human judgment fits in. We sat down with three experts at JMP to discuss the evolving landscape of predictive modeling. Ryan Lekivetz, Director of Advanced Analytics R&D; Chris Gotwalt, Chief Data Scientist; and Russ Wolfinger, Distinguished Research Fellow, each bring a unique perspective on model development and why human judgment remains essential to building models that deliver meaningful insights.
Designing the data, not just modeling it
What trends in predictive modeling are you most excited about right now? Why?
Lekivetz: What excites me most is the shift from passively modeling observed data to actively designing the data we learn from, what I think of as “Big DOE.”
Coming from a design of experiments perspective, I am used to thinking about problems before the data exists, which means choosing what to run, where to sample, and how to learn as efficiently as possible. What is changing now is that we can apply that thinking at much larger scales. In many modern settings, we have at least partial control over inputs and the ability to iterate quickly. DOE is no longer limited to small, carefully controlled studies. It becomes part of an ongoing, adaptive loop.
The result is that predictive modeling moves beyond curve fitting and becomes guided exploration of a system. When we do this well, we are not just building better models. We are building better data, on purpose.
Automation concentrates responsibility, it doesn't remove it
As automation handles more of the mechanical steps in model building, where does human judgment become nonnegotiable? What can scientists and engineers do to remain meaningfully in control rather than just clicking “run”?
Lekivetz: Automation is very good at optimizing a defined objective, searching model space, and executing workflows. But it assumes the objective is correct, the data is representative, and the evaluation is meaningful. Those assumptions are exactly where things break, and where humans must stay in control.
There are three places where this matters most:
- Defining the objective
What are we actually trying to optimize, and does that metric reflect reality? Many failures start here, not in the modeling.
- Designing the data
If we have any control over how data is collected, through experiments or sampling, then we are making design decisions whether we acknowledge it or not. Being explicit about those choices is what separates learning from just accumulating data.
- Validation under the oracle problem
In many settings, we do not observe ground truth in a clean or complete way, which makes it easy to overestimate performance with internal metrics alone, especially when conditions shift or key variables are missing. This situation is where models look right until they are used.
At JMP, a lot of what we do is intentionally keep the human in that loop. The goal is not to remove judgment; it is to surface the right information at the right time so the user can make informed decisions. It means exposing assumptions, showing uncertainty, and making it clear where results depend on choices the user controls.
Staying in control is less about manually executing every step and more about owning the structure of the problem and the credibility of the result. Automation concentrates responsibility. It does not remove it.
Wolfinger: Human judgment is always nonnegotiable. Each of us must build and maintain trust in any tool we utilize and accept full accountability for results we produce.
The real skill is knowing when a model is wrong
There's a real tension between democratizing predictive modeling and ensuring the people doing it actually understand what's happening under the hood. Is lowering the barrier to entry a net positive for the field, or are we quietly building a generation of modelers who can't explain what their models are doing?
Lekivetz: It is both a positive and a real risk, but I think the concern is often framed incorrectly. The main issue is not whether someone can explain the model. It is whether they can recognize when it is wrong, or when the result is more uncertain than it appears.
We have a long history of using powerful methods that people learn to apply effectively without needing to understand every detail. What is different now is the speed, scale, and confidence with which models can be produced and deployed. It reduces the natural friction that used to force deeper scrutiny. The real skill is understanding failure and uncertainty, which includes recognizing when the data does not support the question, when evaluation is overly optimistic, when conditions have shifted, or when the model output is being interpreted with more confidence than it deserves.
Lowering the barrier to entry is a net positive, but only if it comes with visibility into assumptions and uncertainty, and a clear expectation that results should be questioned, not just accepted. The goal is not that everyone can fully explain every model, but that they understand when to trust it, when to doubt it, and how to investigate further.
The cost of skipping the slow work
In a world where a model can be built in minutes, what's lost when teams skip the slow work: exploratory analysis, assumption checking, truly understanding the data-generating process? Is the culture of “just ship a model” changing how organizations think about statistical risk?
Lekivetz: Exploratory analysis and assumption checking are how we figure out what is actually going on in the data. They help us see where we have good coverage and where we are stretching the model beyond what the data can support. If you skip that step, it becomes easy to build models that look accurate but do not reflect how the system really behaves.
The “just ship a model” mindset puts too much weight on performance metrics after the fact. In practice, a lot of that risk is already baked in by how the data was collected and whether it reflects the decisions the model is supposed to support.
As models get easier to build, the pressure is to move faster. The teams that handle this well are the ones that know when to slow down and invest in understanding and improving their data before trusting the result.
Gotwalt: There is a big difference between the use of statistics for carefully proving something, like whether a new drug is effective, and fitting models that are going to be used predictively. When you are proving something, there is a lot more need for understanding the data-generating mechanism and checking assumptions than when you are predicting something. A lot of the time, it would be overkill to apply the same standards in both situations. Using models that have built-in robustness to missing data and outliers, like tree-based methods, has become popular since they worked well with messy data out of the box and one could quickly find an accurate model with them. Different situations present different risks, and people will generally do the easiest thing possible until problems arise that force them to do otherwise.
Wolfinger: The importance of data prep, cleaning, exploration, and feature engineering is emphasized highly and appreciated greatly in both cultures. This is common ground we can certainly build upon.
The tools are ahead of the discipline needed to question them
Whose responsibility is it to actively identify potential bias and flag dubious modeling choices: the data scientist, the organization, or the tools themselves? Is the field doing enough on that front?
Lekivetz: Ultimately, it sits with the person using the model. Tools can help and organizations can set expectations, but someone still has to decide whether the result makes sense and whether it should be trusted. Part of that is mindset. You have to go in assuming there could be something wrong, whether that is in the data, the setup, or the tool itself. A lot of problems show up when people treat outputs as more reliable than they really are. That kind of mindset is shaped by the organization. Culture and expectations matter, especially in how much results are questioned versus accepted.
We are getting better tools, but I do not think the field has fully caught up on the human side. The ability to produce results has moved faster than the discipline needed to question them.
Gotwalt: Ultimately it will be up to the organization to define policy that sets expectations for how people will responsibly use these tools. As it is, I do not think enough is being done on this and other fronts. We are still in the Wild West, the irrationally exuberant days, while the tools are rapidly improving in ways unimaginable a couple years ago. Things will settle down eventually and policy will be defined in response to situations that arise where people went too far and bad things happened, such as dubious models being deployed without sufficient oversight.
The tools will keep getting faster. The question these three keep circling back to is whether the people using them are getting more careful, not just more capable.
Interested in learning more?
In this on-demand talk, Russ Wolfinger walks through how statistics and modern AI actually work together, applications with image, text, and tabular data, and where domain expertise still does the heavy lifting.