Web Analytics
Prachi Singh

October 6, 2026

How to Take a Machine Learning Model From Notebook to Production in 2027

In the third quarter of 2021, Zillow bought roughly 9,700 homes. Its home-buying arm, Zillow Offers, used pricing algorithms to estimate what each house was worth and how much to pay for it. A few weeks later the company recorded a $304.4 million inventory write-down, because it had paid more for those homes than it now expected to sell them for. On November 2, 2021, Zillow's board decided to wind the business down, a decision the company said would cut about a quarter of its workforce. All of this is in Zillow's own SEC filings.

CEO Rich Barton said home prices had proved far harder to forecast than expected. The filing also pointed to capacity limits and an unusual pandemic-era housing market.

People often file this under "AI gone wrong." Look closer and most of the damage happened after the model left the lab. The market moved, the system kept buying at volume, and small pricing errors got multiplied by thousands of real purchases. Getting a machine learning model to production is mostly about surviving that second part.

What follows is the full route, from the notebook a data scientist hands over to a system that answers real requests and warns you when it drifts. Technical terms are explained as they come up.

The numbers behind the notebook problem

54%

Average share of AI projects that moved from pilot to production in a Gartner survey of 699 organizations (run October to December 2021, published August 2022). Gartner's 2019 survey put it at 53%.

Nearly 2 in 3

Organizations that have not yet begun scaling AI across the enterprise, according to McKinsey's State of AI 2025 survey of 1,993 participants in 105 countries.

39%

Respondents in that same McKinsey survey who report any impact on enterprise-level EBIT (operating profit) from AI.

88%

Organizations using AI regularly in at least one business function in 2025, up from 78% in 2024 (McKinsey).

2 Dec 2027

New start date for EU AI Act rules on stand-alone high-risk AI systems, under the provisional Digital Omnibus deal reached on May 7, 2026. The original date was August 2, 2026.

$1.7B to $2.98B

Range of recent MLOps market size estimates. The research firms disagree, as explained below.

Those market estimates do not match. Global Market Insights valued the MLOps market at USD 1.7 billion in 2024. SkyQuest put 2024 at USD 1.97 billion. Fortune Business Insights valued 2025 at USD 2.98 billion and projects much faster growth than the others. Each firm defines the market differently (some count consulting services, some only software), so read them as a sign of direction rather than a precise size.

The Gartner and McKinsey figures point the same way. Building a model is no longer the hard part for most companies. Keeping one running and trusted is.

What "production" actually means

A notebook, usually a Jupyter notebook, is a document where you write code in small blocks, run each block, and see the result right underneath. It is great for exploring ideas and fragile for anything a business depends on. Cells can run out of order, the data is often a CSV file downloaded months ago, and the whole thing lives on one laptop.

A model is "in production" when it is part of a real product or business process and nobody has to babysit it. In practice that means four things are true:

▪         It runs without its author being present or awake.

▪         It receives data the same way every time, from systems that are documented.

▪         It can be updated or rolled back to an older version in minutes.

▪         Someone gets told, automatically, when it starts behaving strangely.

A home cook who makes a great biryani has proved the recipe works. A restaurant serving 400 plates of it on a Saturday night has a different job: the same taste every time, ingredients on schedule, and a backup plan when the gas runs out. The notebook is the home kitchen. Production is the restaurant.

Why so many models get stuck halfway

In 2015, a group of Google engineers led by D. Sculley published a paper at the NeurIPS conference called "Hidden Technical Debt in Machine Learning Systems." Its best-known diagram shows the machine learning code as a tiny box inside a much larger system of data collection, data checks, feature extraction, configuration, serving, and monitoring. The model is a small fraction of what has to be built and maintained.

That gap explains most stalled projects:

▪       The notebook only works on the author's machine, with library versions nobody wrote down.

▪       Training used a clean historical snapshot, while live data arrives late, half-empty, or in a new format.

▪       Nobody owns the model after launch, so problems sit unnoticed until a customer complains.

▪      The target was a model score, and no one agreed with the business on what "good enough" means in money or time saved.

▪      Monitoring was planned for "phase two," and phase two never got funded.

Gartner's Frances Karamouzis made a related point when the 2022 survey came out: organizations still struggle to tie the algorithms they build to a clear business value, which makes it hard to justify the money needed to run them properly. Shipping a model is partly engineering and partly an argument you have to win inside the company.

The route from notebook to production, stage by stage

Each stage below has a clear output. If you cannot point to that output, the stage is not finished.

Stage 1: Agree on the job and the number

Before touching code, write down what decision the model supports, who acts on it, and what a wrong answer costs. A fraud model that blocks a genuine customer loses a sale. One that misses a fraudster costs the transaction plus fees. Those costs are rarely equal, and the model should be tuned with that in mind.

Pick two numbers. The first is a model metric, such as precision (of the cases the model flagged, how many were truly fraud) or recall (of all the real fraud, how much the model caught). The second is a business metric, such as chargebacks per thousand orders. Then build a baseline, a simple rule like "flag large orders from new accounts." If the model cannot clearly beat that rule, stop before spending months on deployment.

Output: a one-page brief with the decision, the costs of each error type, both metrics, and the baseline result.

Stage 2: Turn notebook cells into real code

Move the logic out of notebook cells and into plain Python files with named functions. Put everything in Git, a version control tool that records every change and who made it. Pin library versions, meaning you record the exact release of each package so the code behaves the same next year. Fix random seeds so training gives repeatable results.

Add tests, both ordinary ones ("does this function handle an empty list?") and data checks ("are all ages between 0 and 120?"). The goal is a training pipeline that runs with one command on a clean machine.

Skill gaps show up here. Companies that hire AI/ML developers based on polished notebook demos sometimes find the person has never packaged code for anyone else to run. Ask candidates for a project someone else has deployed.

Output: a repository where a new team member can retrain the model by following the README.

Stage 3: Version the data and the model together

Code versioning is not enough, because the same code trained on different data gives a different model. Use a model registry, which is a catalog that stores each trained model with a version number, the exact data it learned from, its test scores, and who approved it. MLflow is a widely used open source option. DVC handles data versioning. The big cloud platforms (Amazon SageMaker, Google Vertex AI, and Azure Machine Learning) each include their own registry.

The test for this stage: if a customer disputes a decision made last Tuesday, can you name the model version behind it and the data that trained it within a few minutes? 

Stage 4: Package it so it runs anywhere

A container, most often built with Docker, is like a sealed lunchbox holding the model, the code, and every library it needs. The same box runs on a laptop, a test server, and the cloud, which removes the classic "it worked on my machine" problem.

Inside the container, the model is usually wrapped in an API, a doorway that other software can knock on with a request and get a prediction back. Small teams often write this with FastAPI. Larger ones use serving tools such as BentoML, KServe, or NVIDIA Triton, or a managed endpoint from their cloud provider. Most practical choices in machine learning model deployment get made here and in the next stage.

Output: a container image that starts, loads the model, and answers a test request correctly.

Stage 5: Choose how predictions reach people

Not every model needs to answer in milliseconds. The four common serving patterns compare like this:

Serving pattern

How it works

Typical speed

Good fit

Watch out for

Batch

Runs on a schedule over many records at once and saves the results

Minutes to hours

Nightly churn scores, weekly demand forecasts, monthly risk reports

Predictions can be stale by the time someone uses them

Real-time API

Answers each request the moment it arrives

Milliseconds to about a second

Fraud checks at checkout, search ranking, product recommendations

Speed limits, and the cost of servers that stay on all day

Streaming

Reacts to a continuous flow of events from a tool like Apache Kafka

Seconds

Sensor alerts, live delivery time estimates, account takeover detection

Harder to build, test, and debug

On-device

The model runs on the phone, camera, or machine itself

Very fast, works offline

Keyboard suggestions, factory camera inspection, voice wake words

Updating models spread across thousands of devices

If the business can live with daily predictions, start with batch. It is cheaper and far more forgiving. Many teams build real-time systems because they sound modern, then learn the scores only get read each morning.

Stage 6: Release in small, reversible steps

Never swap models in a single move. Safer options:

1.      Shadow mode: the new model runs next to the current one and its predictions are logged but not used. You compare the two on real traffic with zero risk to customers.

2.      Canary release: a small slice of traffic, say 5%, goes to the new model while you watch error rates and business numbers.

3.      A/B test: traffic is split between versions long enough to measure a real difference in outcomes.

Keep the previous model in the registry and make rollback a single command or a single switch in a settings file. Practice it once before you need it.

Stage 7: Watch, learn, and retrain

Here the MLOps vs traditional DevOps split becomes obvious. Ordinary software stays correct until someone changes the code. A model can get worse while the code sits untouched, because the world it describes has shifted. Retraining can follow a schedule, trigger when input data shifts past a threshold, or trigger when accuracy drops. Many teams combine a schedule with a drift alarm.

MLOps vs traditional DevOps: where the differences show up

DevOps is a set of habits and tools that let software teams ship code often and safely. Its core is CI/CD (continuous integration and continuous delivery), where every change is tested automatically and released through a repeatable pipeline. Comparing MLOps vs traditional DevOps comes down to this: MLOps keeps all of that and adds data and models to the list of things that can break.

Area

Traditional DevOps

MLOps

What changes system behavior

Code and configuration

Code, configuration, training data, and model weights

What gets versioned

Source code and settings

Code, data snapshots, feature definitions, models, and training settings

What tests check

Does the function return the expected output?

That, plus: is accuracy above the agreed line, does the data look as expected, are results fair across user groups?

How failure looks

Crashes, error pages, slow responses

Often silent: the system keeps answering, but the answers get worse

When updates are needed

When someone ships a change

Also when the world changes, even with no new code

Who is involved

Developers and operations engineers

Also data scientists, data engineers, and domain experts

Rolling back

Redeploy the previous build

Redeploy the previous model and confirm the feature pipeline still matches it

The failure row causes the most pain. A crashed web server sets off alarms within seconds. A credit model that slowly starts approving riskier applicants sets off nothing, because technically it is working fine.

ML model monitoring in production: what to watch

Good ML model monitoring in production works in layers, from the plumbing up to the business result. Skipping a layer leaves a blind spot.

▪      Service health: is the endpoint up, how long do responses take, and how many requests fail?

▪      Input data: are missing values rising, are numbers outside their usual range, have new categories appeared? A shift in inputs compared with training data is called data drift.

▪      Predictions: has the mix of outputs changed? If the share of orders flagged as risky jumps from 2% to 9% overnight, something upstream probably broke.

▪      Real outcomes: once true answers arrive, how accurate was the model? When the link between inputs and outcomes changes, such as what drives home prices in 2020 compared with 2022, that is concept drift, and retraining is the usual fix.

▪      Business result: are chargebacks, sales, or support tickets moving the way the brief in Stage 1 said they should?

Real outcomes are often delayed. You learn whether a loan defaults months after approving it, so watch early proxies like missed first payments. Common drift checks include the Population Stability Index, which scores how far a distribution has moved. Open source tools like Evidently AI, and services like SageMaker Model Monitor and Vertex AI Model Monitoring, run these checks for you. 

The hard parts nobody shows in the demo

Demos run on clean data, polite users, and quiet servers. Production offers none of those.

When data goes missing

A feature the model needs may be missing when a request arrives. An upstream service is down, a new field is empty for older customers, or a partner changed a file format. The usual responses:

▪      Fill a default value. Quick, but filling with an average can hide real problems and push predictions toward the middle.

▪      Mark it as "unknown" and train the model with that category, so it learns how to behave without the value.

▪      Fall back to a simpler model trained without that feature.

▪      Decline to score and route the case to a rule or a person.

Missing values can carry meaning. Applicants who leave income blank may behave differently from those who fill it in. Watch for the quieter version too. If training computed "days since last purchase" using full history but the live system only stores 90 days, every long-dormant customer looks like a 90-day customer. Nothing breaks, and the predictions are wrong.

When signals disagree

Real systems rarely have one source of truth. The model gives a low fraud score, but a merchant rule says block. Or years of clean payment history clash with a brand-new device in another country.

Write the decision policy down before launch: which source wins, in what order, and at what confidence. A common pattern uses three zones, with a grey middle zone that goes to human review or a one-time password check. Log every disagreement with its eventual outcome. Those cases are the best training material you will get, because they show exactly where the model and the rules part ways.

When the answer is needed in milliseconds

Real-time decisions run on a time budget. If checkout must load in a fraction of a second, the model gets only a slice after the network, database, and payment provider take theirs. That shapes the design:

▪      Precompute heavy features and keep them in a feature store, a shared database that training and live serving both read from.

▪      Set a hard timeout on the model call and decide in advance what happens when it fires (approve, decline, or use a cached score).

▪      Use a smaller, faster model when the large one is too slow. A technique called distillation trains a small model to copy a big one's answers.

When the input is something the model never saw

Edge cases are where trust gets lost. New users have no history (the cold start problem). Someone types 212 into an age field. A festival or a viral post creates a week unlike anything in the training data. Some users try to game the system.

Put a validation layer in front of the model to reject impossible inputs. Let the model abstain when confidence is low. Keep a growing test set of strange past cases and rerun it before every release. Good ML model monitoring in production reports these cases as their own stream rather than burying them inside an overall accuracy average, where a 0.5% slice of disasters looks like noise.

When traffic spikes or the system grows

A sale day can bring several times the usual traffic. Autoscaling, where the platform adds servers as load rises, helps, but GPU servers with large models can take minutes to start. Teams keep a few warm spares, group requests into small batches, and put a queue in front of the model so bursts wait instead of crashing it.

Plan a graceful degradation order in advance: cached scores first, then a lighter model, then a simple rule, and failure only as a last resort. Customers rarely notice a slightly worse recommendation. They always notice an error page.

Scale also creates feedback loops. When a fraud model blocks a transaction, you never learn whether it was really fraud. A recommender showing only popular items makes them more popular. A common fix is a small random holdout of traffic where the model's decision is not applied, so honest data keeps coming in.

Finally, scale means many models. Gartner's 2022 survey found that 40% of organizations reported having thousands of AI models deployed. At that size, machine learning model deployment has to become a routine with shared templates, standard monitoring, and a named owner for every model.

Build it yourself, hire, or bring in a partner

There are three realistic routes, and many companies mix them. Your existing team can do it if you already have data engineers and backend developers comfortable with cloud infrastructure.

You can hire AI/ML developers when machine learning is central to the product and you expect to run several models for years. Look for people who can show:

▪         Python code that others have run, with tests and clear documentation

▪         Comfort with SQL and data pipelines, since most model problems start in the data

▪         Experience with containers and at least one major cloud platform

▪         A model they personally kept healthy after launch, including a drift problem they caught

▪         The ability to explain a trade-off, like speed against accuracy, to non-technical colleagues 

You can bring in a partner offering machine learning development services when you need a first model live quickly or want an outside team to set up pipelines your people will later run. Before signing, ask who owns monitoring after launch, how retraining is paid for, whether they can demonstrate a rollback, and what you keep at the end.

In Gartner's survey, 72% of executives said they had, or could find, the AI talent they needed. Talent was not the main barrier. Tying models to business value was. Whichever route you choose, insist on the Stage 1 brief first.

What changes in 2027

The first is regulation. Under the provisional Digital Omnibus agreement reached on May 7, 2026, EU AI Act obligations for stand-alone high-risk AI systems now apply from December 2, 2027, and obligations for AI built into regulated products, such as medical devices and machinery, apply from August 2, 2028. Stand-alone high-risk uses include hiring and HR tools, credit scoring, access to education, critical infrastructure, and law enforcement. The Act asks providers of these systems for risk management, data governance, record keeping, and human oversight. A model registry, versioned data, and monitoring logs are the records regulators will want, so the habits in this guide double as compliance groundwork. Good providers of machine learning development services ask about EU exposure early, since systems used on people in the EU can fall in scope wherever the company is based.

The second shift is that more "models" are now large language model features such as chatbots, document summaries, and AI agents. McKinsey's 2025 survey found that 62% of organizations were at least experimenting with AI agents and 23% were scaling an agentic system somewhere in the business. These need the same discipline. Version prompts like code. Keep an evaluation set of real questions with known good answers and rerun it before each change. Track cost per request, and watch for output drift when the provider updates its model.

A pre-launch checklist

Before you push a machine learning model to production, confirm that:

▪         The Stage 1 brief exists and the business owner has signed off on it

▪         Training runs with one command on a clean machine

▪         Every model version in the registry links to its data, code, and test scores

▪         The serving pattern matches how quickly people actually need answers

▪         Input validation rejects impossible values

▪         A fallback exists for missing data, timeouts, and overload

▪         The decision policy for conflicting signals is written down

▪         Rollback has been tested at least once

▪         Alerts have named owners and a written first step 

Where to start this week

Pick the model sitting in a notebook that people keep calling "almost ready." Write its Stage 1 brief and test it against a simple rule. If it wins, move the code into a repository and log its inputs and outputs before building anything fancy. Those logs will tell you more than any notebook score.

Zillow's model did not fail in a notebook. It failed in a market that moved faster than the system around it could react. The teams that do well in 2027 will be the ones that build that system first and treat the model as one replaceable part of it.

Prachi Singh

With a love for connecting with people and a flair for communication, Prachi's expertise in digital marketing is unmatched. Her strategic approach to campaigns ensures our brand's story reaches far and wide, making an impact on the lives of countless individuals.

Frequently Asked Questions

It depends more on the data and the stakes than on the model. A batch model on clean, documented data can go live in weeks. A real-time system in lending or hiring usually takes months because of testing, review, and documentation. Moving a machine learning model to production slows down most when teams skip Stage 1 and argue about targets after launch.

No. The MLOps vs traditional DevOps question is not a choice between two camps. MLOps sits on top of DevOps. You still need version control, automated testing, and repeatable releases. MLOps adds data versioning, model registries, drift monitoring, and retraining to that base.

At a minimum, track uptime and response time, log every input and prediction, watch the share of each prediction type over time, and compare predictions against real outcomes whenever those arrive. That basic level of ML model monitoring in production catches most serious problems early and costs very little to set up.

If machine learning is the core of your product, plan to hire AI/ML developers eventually, since that knowledge should live inside the company. For a first model, or when you need something live before you can hire, machine learning development services can get you there faster. Make sure the contract covers handover, documentation, and ownership of everything built.

Usually not at the start. Many classic models, such as gradient boosted trees for scoring and forecasting, run fine on ordinary CPU servers. Kubernetes and GPUs make sense once you run many models, serve large neural networks, or face heavy real-time traffic. For a first machine learning model deployment, a single container on a managed cloud service is often enough.

  • Hourly
  • $20

  • Includes
  • Duration: Hourly Basis
  • Communication: Phone, Skype, Slack, Chat, Email
  • Project Trackers: Daily reports, Basecamp, Jira, Redmi
  • Methodology: Agile