• Key takeaways

    • MIT’s NANDA research found that 95% of enterprise generative AI pilots fail to deliver measurable ROI, despite tens of billions of dollars in enterprise investment — a number that’s held up across multiple independent 2026 studies from IDC, RAND, and S&P Global.
    • The gap isn’t about model quality. It’s architectural and organizational: unready data infrastructure, infrastructure costs that run 3–5x initial projections, no evaluation or drift monitoring, and unclear ownership when something breaks.
    • Budget allocation makes the divide worse: over half of 2025 AI budgets went to high-visibility sales and marketing pilots, while the real returns showed up in less glamorous back-office automation.
    • The organizations in the successful 5% share a pattern: tightly scoped use cases, agents connected to real institutional data rather than a chatbot with a system prompt, and a blend of internal AI specialists with outside expertise — which one industry study put at a 67% success rate versus 22% for IT-only builds.

    The Number That Should Stop Every AI Steering Committee

    MIT’s Project NANDA published a study in 2025 — “The GenAI Divide: State of AI in Business” — that’s become the reference point for nearly every serious conversation about enterprise AI ROI in 2026. The finding: 95% of organizations see no measurable return to the income statement from their generative AI pilots. Only a small minority — around 5% — extract real value at scale. Global AI spending is projected to hit $2.5 trillion in 2026, yet only about 6% of companies report AI contributing more than 5% to EBIT. That gap between spend and impact is what MIT calls the GenAI Divide, and it isn’t closing on its own.

    It’s worth being precise about what this number does and doesn’t mean. It’s not that generative AI doesn’t work. Individual employees are getting real value out of tools like ChatGPT and Copilot — MIT’s research notes over 80% of organizations have piloted these tools and roughly 40% report some level of deployment. The problem is that those individual productivity gains almost never survive contact with an organization’s budget process, procurement, data infrastructure, and change management — which is exactly where ROI actually has to get proven to a board.

    Where the Money Actually Goes Wrong

    A few root causes show up consistently across the 2026 research:

    Infrastructure cost surprise. Production GenAI deployments typically run 3 to 5 times the initial cost projection, which quietly kills the ROI case before anyone officially declares the pilot a failure.

    Data that was never production-ready. An agent is only as good as the institutional data it’s grounded in. Projects scoped around a flashy demo, running on synthetic or narrow test data, look impressive in a steering committee meeting and collapse the moment they hit real production traffic.

    No evals, no drift monitoring. Teams that can’t measure whether a model change improved or regressed quality have no way to defend a scaling decision, let alone an ROI claim.

    Misallocated budget. Over 50% of 2025 AI budgets went to sales and marketing pilots — high visibility, comparatively low ROI — while the real returns showed up in less exciting back-office automation. S&P Global’s 2025 Voice of the Enterprise survey found 42% of companies abandoned most of their AI initiatives that year, up sharply from 17% the year before, scrapping nearly half their AI proofs of concept before they ever reached production.

    What the Successful 5% Actually Do

    The pattern in the winning minority isn’t exotic. It’s disciplined. Successful organizations scope tightly around a single, measurable business outcome instead of a broad transformation narrative. They connect agents to real institutional data and existing systems, rather than shipping a chatbot wrapped around a system prompt. They instrument evaluation and drift monitoring from day one instead of bolting it on after launch. And they blend internal AI specialists with outside expertise — one industry analysis found pilots built this way hit a 67% success rate, compared with 22% for internal-IT-only builds, because most organizations don’t yet have the combination of AI engineering maturity and domain judgment in-house to do it alone.

    Where This Connects to the Boring Infrastructure Argument

    This is the part that doesn’t make it into the flashy AI headlines, and it’s the argument I find myself making most often: the GenAI ROI gap and the data architecture gap are the same problem wearing different clothes. An AI agent making decisions off a fragmented, ungoverned data estate isn’t an AI problem waiting on a better model — it’s a data foundation problem that a better model can’t fix. I’ve watched this play out directly: a cloud data platform modernization I led earlier in my career had no AI story attached to it at all, and it became the reference architecture the broader practice later pointed to when pitching AI-driven work to other clients. The unglamorous infrastructure work is usually the actual prerequisite for the AI story people want to tell.

    Healthy skepticism about GenAI ROI isn’t anti-AI. It’s the discipline that separates the 5% from the 95% — and it starts with treating “is our data ready for this” as the first question, not the one you answer after the pilot already has executive buy-in.

    Frequently Asked Questions

    Why do most enterprise AI pilots fail?

    Primarily for organizational and architectural reasons, not model quality: unready production data, infrastructure costs running 3–5x initial projections, absent evaluation and drift monitoring, and unclear ownership once something breaks in production.

    What is the real ROI of generative AI in 2026?

    Highly bimodal. MIT’s research found roughly 95% of enterprise GenAI pilots show no measurable financial return, while the successful 5% report returns as high as $3.70 for every dollar spent — a gap driven by execution discipline, not access to better models.

    How do you measure GenAI ROI properly?

    Define the specific business outcome and baseline metric before building, instrument evaluations and drift monitoring from day one, test at production-shaped traffic rather than a small synthetic sample, and assign clear operational ownership before launch rather than after an incident.

    Is generative AI overhyped?

    The technology itself isn’t the bottleneck — the surrounding system is. Individual productivity gains from tools like ChatGPT and Copilot are real and widely reported; what’s overhyped is the assumption that those gains translate automatically into enterprise-level P&L impact without the underlying data and governance work.


    Sources referenced


    About the author: Abhishek Srivastava is a Senior Data & AI Strategy executive with 20+ years across enterprise data architecture, cloud platforms, and AI/ML delivery — with a track record spanning both consulting-side AI deal strategy and in-house delivery accountability. He’s open to senior data & AI leadership conversations. Connect on LinkedIn to continue the conversation.

  • Key takeaways

    • AI governance stopped being optional in 2026. The EU AI Act is in full enforcement, NIST’s AI Risk Management Framework has become the de facto US enterprise standard, and Gartner estimates organizations without a formal program face materially higher rates of AI-related incidents.
    • Three frameworks anchor most enterprise programs: NIST AI RMF (voluntary, US-centric, built around Govern/Map/Measure/Manage), the EU AI Act (binding, risk-tiered, applies to any org whose AI outputs reach EU users), and ISO/IEC 42001 (the first certifiable AI management system standard).
    • The newest gap: every one of those frameworks assumes a human in the loop. Autonomous agents that connect to tools and act on their own — the direction most enterprise AI is heading — remove that assumption, and governance programs built in 2023–2025 haven’t caught up.
    • The practical starting point isn’t picking one framework. It’s building a living AI system registry, classifying each system against risk tiers, and assigning a named owner — because you can’t govern what you haven’t inventoried.

    Why This Became Urgent, Not Optional

    For most of the last decade, “AI governance” was a slide in a strategy deck. That changed in 2026. The EU AI Act, which entered force in August 2024, is now in full application, and it applies to any organization whose AI systems affect EU users — regardless of where that organization is headquartered. In the US, the NIST AI Risk Management Framework remains technically voluntary, but it has become the reference architecture enterprise customers and federal procurement processes actually check for. And the operating cost of skipping governance doesn’t just show up as regulatory fines — it shows up earlier, as shadow AI deployments, agent sprawl, and incidents that surface during a board review instead of a compliance audit.

    The Three-Framework Stack

    Most global enterprises aren’t choosing one framework — they’re layering two or three, and the good news is they share a common backbone.

    NIST AI RMF organizes around four functions — Govern, Map, Measure, Manage — that have become a shared vocabulary across the other frameworks. It’s voluntary, but it’s foundational: if you build your internal program around these four functions, mapping obligations across other jurisdictions gets far easier.

    The EU AI Act is the first binding, horizontal AI regulation — meaning it cuts across industries rather than targeting one sector. It classifies systems into four risk tiers (unacceptable, high, limited, minimal). Annex III defines the high-risk categories, and Article 26 places specific obligations on deployers: conformity assessments, human oversight mechanisms, and incident reporting. Those obligations are already active for most high-risk categories.

    ISO/IEC 42001 adds a certification path — the first international standard for an AI Management System. Its structure deliberately mirrors ISO 27001, which makes it a natural extension for any organization that already has an information security certification program in place.

    A sensible rollout: NIST AI RMF for risk methodology, ISO 42001 for certifiable infrastructure, EU AI Act obligations layered in for EU-facing operations. Industry estimates put full implementation of all three at roughly 8–12 months for a moderately complex organization — largely because the frameworks share so much common ground in risk assessment, human oversight, and documentation requirements that you effectively build the underlying program once.

    The Gap Nobody’s Framework Covers Yet

    Here’s the part I find most relevant to where enterprise AI is actually headed: every major framework — NIST AI RMF, ISO 42001, the EU AI Act — quietly assumes a human is in the loop, making the decision or approving the action. Agentic AI, where a model connects to external tools and data sources through something like the Model Context Protocol and acts on its own, removes that human from the chain entirely. NIST’s Govern function assumes an accountable human decision-maker. The EU AI Act’s human-oversight obligations assume someone to inform and someone to oversee. ISO 42001’s controls assume a human-run process.

    Singapore’s Infocomm Media Development Authority moved first here, launching a governance framework specifically for autonomous agents in January 2026. In the US, individual states aren’t waiting either — Texas’s Responsible Artificial Intelligence Governance Act took effect January 1, 2026. Expect more state-level and sector-specific agentic-AI addenda to existing frameworks over the next 12–18 months, because the gap is real and the technology isn’t slowing down for the frameworks to catch up.

    What I’d Actually Tell a Leadership Team to Do First

    Skip the temptation to pick a framework before you know what you’re governing. Build a living AI system registry first — every model, every API integration, every vendor-embedded AI capability in active use. Classify each one against the EU AI Act’s risk tiers even if you’re US-only; it’s the clearest risk vocabulary available and it forces the conversation. Assign a named owner to each system, not a committee. And build your review categories around the actual domains that trigger real risk — data privacy and confidentiality, ethical and bias review, architecture and security, and cyber/data-residency — rather than a generic “AI ethics checklist” that nobody actually applies at decision time. That’s the structure I’ve used running governance reviews with clients directly, and it holds up better than a framework binder nobody opens until an auditor asks for it.

    Frequently Asked Questions

    Is the NIST AI RMF mandatory?

    No — it’s voluntary. But it’s increasingly referenced in federal procurement requirements and used by enterprise customers as a baseline for vendor due diligence, which makes it a practical necessity for most B2B and enterprise-facing organizations even without a legal mandate.

    Does the EU AI Act apply to US companies?

    Yes, if their AI systems are deployed to or affect users in the EU, regardless of where the company is headquartered. This is the same extraterritorial logic GDPR used, and it’s already active for high-risk system categories.

    What does ISO 42001 certification actually prove?

    It demonstrates to customers, regulators, and partners that an organization’s AI management system meets an audited international standard — similar to what ISO 27001 does for information security. It’s a certification path, not a legal requirement, but it’s increasingly showing up as a procurement requirement in enterprise deals.

    Do existing AI governance frameworks cover autonomous AI agents?

    Not fully. NIST AI RMF, the EU AI Act, and ISO 42001 were all built assuming a human is in the loop making or approving decisions. Autonomous agents that act through tool integrations without human sign-off fall into a governance gap that only a handful of jurisdictions — Singapore and Texas among the first — have started addressing directly.


    Sources referenced


    About the author: Abhishek Srivastava is a Senior Data & AI Strategy executive with 20+ years across enterprise data architecture, cloud platforms, and AI/ML governance — with hands-on experience navigating privacy, ethics, and architecture review processes on real enterprise AI deployments, from both the consulting and in-house sides. He’s open to senior data & AI leadership conversations. Connect on LinkedIn to continue the conversation.

  • Key takeaways

    • Medallion architecture (Bronze → Silver → Gold) is still the default pattern for organizing a data lakehouse, and for good reason — it turns a data lake into something trustworthy instead of a swamp.
    • Its structural weakness is latency: the multi-hop refinement process that makes it great for BI and reporting adds propagation delay that breaks real-time, agentic AI decisioning.
    • Its second weakness is organizational, not technical: most teams treat it as a storage convention when it’s really a team contract — an agreement about who owns data quality at each layer.
    • The fix isn’t abandoning medallion. It’s being deliberate about where it applies, adding a semantic layer on top, and building the parts of your data estate that need sub-second freshness outside the medallion flow entirely.

    Why Medallion Became the Default

    Databricks popularized the Bronze/Silver/Gold pattern around 2019–2020 while promoting the lakehouse paradigm, though the underlying idea — progressively refining data through distinct stages — has roots going back to classic data warehousing. It solved a real problem: data lakes had promised unlimited flexibility and instead produced data swamps — millions of files with no lineage, duplicated transformations, and dashboards nobody trusted.

    The pattern is simple to describe. Bronze lands raw data exactly as it arrives, with no business logic, so it stays complete and replayable. Silver cleans, deduplicates, and conforms it into a shared, trustworthy foundation. Gold aggregates it into business-ready tables for BI dashboards, ML features, or executive reporting. It’s vendor-agnostic — you can run it on Databricks, Snowflake, Microsoft Fabric, or BigQuery — which is part of why it spread so widely.

    Where It’s Actually Cracking

    I’ve spent two decades building and living with these platforms — inside enterprises and on the consulting side advising them — and the cracks showing up in 2026 aren’t hypothetical. They fall into two categories.

    1. It wasn’t built for real-time, agentic decisions. Medallion’s multi-hop flow — bronze to silver to gold, each a separate processing stage — adds propagation delay. That’s a fine tradeoff for a dashboard refreshed nightly. It’s a structural limitation when an AI agent needs a decision inside a tight validity window. As AI workloads shift from “generate a report” to “take an action right now,” a growing body of practitioner analysis is calling this out explicitly: the pattern excels at analytics and BI, and breaks down for real-time automated decisioning.

    2. It’s a team contract wearing a storage diagram’s clothes. The most common implementation failure isn’t a schema mistake — it’s teams getting the layer names right and the ownership model wrong. Medallion is fundamentally an agreement about who’s accountable for data quality at each stage and what guarantees downstream consumers can rely on. Skip that conversation, and you get exactly what critics of the pattern describe: a heap of constantly-generated bronze data that downstream teams have to sift through themselves, duplicated transformation logic scattered across teams, and rising storage and compute costs with no corresponding increase in trust.

    My Take, From Both Sides of the Table

    I’ve built lakehouse platforms in-house — making the real call on Snowflake, GCP, and CDP tooling when the vendor pitch didn’t match what my team needed day to day — and I’ve advised clients on the same decisions from the consulting side. The pattern I see repeatedly: organizations adopt medallion because it’s the “correct” answer on a whiteboard, then skip the two decisions that actually determine whether it works.

    First, treat the lakehouse as a product, not infrastructure. That means someone owns Silver’s definition of “clean” the way a product manager owns a roadmap — not as a side task bolted onto a data engineer’s sprint.

    Second, don’t force everything through all three layers. If a use case needs sub-second freshness for an agent to act on, build it as its own path — a streaming or serving layer that sits alongside medallion, not underneath it. Medallion inside a lakehouse is still extremely common and still the right default for analytics and ML training data. It’s just no longer the only pattern a mature data platform needs.

    Frequently Asked Questions

    Is medallion architecture dead in the AI era?

    No. It remains the right default for analytics, BI, and most ML training pipelines. It’s the wrong default for real-time, agentic decisioning that can’t tolerate multi-hop propagation delay — that needs a separate, faster path.

    What’s the alternative to medallion for real-time AI?

    Most mature platforms don’t replace medallion outright — they pair it with a streaming or serving layer for time-sensitive use cases, while keeping medallion for historical, analytics, and training-data workloads. Some architects are also pushing toward data-product models with semantic layers to reduce the operational burden medallion places on downstream consumers.

    Do I need medallion architecture if I already have a lakehouse?

    Not automatically. Medallion is a logical pattern, not a required feature of a lakehouse. The right question isn’t “does my lakehouse have medallion layers” — it’s “does my organization have a clear, owned contract for data quality at each stage of refinement.” If the answer is no, adding medallion labels won’t fix it.


    Sources referenced


    About the author: Abhishek Srivastava is a Senior Data & AI Strategy executive with 20+ years across enterprise data architecture, cloud platforms (Snowflake, GCP, Kubernetes), and AI/ML strategy — with experience on both the consulting side (Big 4) and the in-house practitioner side (Fortune 500 retail and travel). He’s open to senior data & AI leadership conversations. Connect on LinkedIn to continue the conversation.

  • Machine learning (ML) has rapidly become a key component of modern businesses, with organizations using ML algorithms to derive insights, automate processes, and improve decision-making. However, deploying and managing machine learning models in production is often a challenging task that requires close collaboration between data scientists and IT operations teams. This is where MLOps comes in – a set of practices and technologies that aim to streamline the ML lifecycle and bridge the gap between machine learning and operations.

    What is MLOps?

    MLOps, short for Machine Learning Operations, is a relatively new term that describes the intersection of machine learning and operations. It encompasses the practices, processes, and technologies used to build, deploy, monitor, and manage machine learning models in production. MLOps borrows from the DevOps culture, which emphasizes collaboration and communication between development and operations teams, as well as the use of automation and continuous integration/continuous delivery (CI/CD) pipelines to streamline software development.

    Why is MLOps important?

    Deploying and managing machine learning models in production is often a complex and time-consuming task that involves multiple stakeholders and steps, such as data preprocessing, model training, validation, deployment, and monitoring. MLOps helps to address some of the key challenges of this process, such as:

    • Versioning and reproducibility: ML models often require specific versions of software libraries and dependencies, and the code used to build them should be versioned and reproducible. MLOps helps to ensure that models are built with the right dependencies and are reproducible, which makes it easier to debug issues and roll back to previous versions if needed.
    • Scalability and performance: Machine learning models can be computationally intensive and require large amounts of data and processing power. MLOps helps to ensure that models can scale and perform well in production by using techniques such as distributed training, load balancing, and resource allocation.
    • Security and compliance: Machine learning models can contain sensitive data or be used in regulated industries, which means that they need to be secured and compliant with relevant regulations. MLOps helps to ensure that models are deployed securely and that they comply with relevant regulations such as GDPR or HIPAA.
    • Monitoring and maintenance: Machine learning models are not static – they need to be constantly monitored and maintained to ensure that they perform well and remain up-to-date. MLOps helps to automate the monitoring and maintenance of models, which frees up data scientists and IT operations teams to focus on more high-value tasks.

    MLOps Practices and Technologies

    MLOps encompasses a wide range of practices and technologies, depending on the specific needs and context of an organization. However, some of the key practices and technologies used in MLOps include:

    • Continuous Integration/Continuous Delivery (CI/CD): CI/CD is a software development practice that emphasizes automation and continuous feedback to improve the speed and quality of software development. In MLOps, CI/CD is used to automate the deployment of machine learning models, as well as to ensure that models are versioned and reproducible.
    • Infrastructure as Code (IaC): IaC is a practice that involves managing infrastructure resources (such as servers or databases) using code. In MLOps, IaC is used to automate the provisioning and management of infrastructure resources that are used to train and deploy machine learning models.
    • Model Versioning: Model versioning is the practice of tracking changes to machine learning models over time. In MLOps, model versioning is used to ensure that models can be reproduced and that issues can be easily tracked and resolved.
    • Automated Testing: Automated testing is the practice of using software tools to automatically test machine learning models. In MLOps, automated
    • testing is used to ensure that models meet certain quality standards and that they perform as expected in different scenarios.
    • Model Monitoring: Model monitoring is the practice of continuously monitoring the performance of machine learning models in production. In MLOps, model monitoring is used to detect and diagnose issues with models, as well as to identify opportunities for improvement.
    • Containerization: Containerization is the practice of packaging software applications (including machine learning models) in lightweight, portable containers that can be run consistently across different environments. In MLOps, containerization is used to simplify the deployment and management of machine learning models, as well as to improve scalability and performance.

    MLOps Workflow

    The MLOps workflow typically consists of the following steps:

    1. Data preparation: In this step, data scientists prepare and preprocess the data that will be used to train the machine learning model.
    2. Model development: In this step, data scientists use machine learning algorithms to train the model on the prepared data. They also evaluate and optimize the model’s performance using techniques such as hyperparameter tuning and cross-validation.
    3. Model deployment: In this step, the trained machine learning model is deployed to a production environment using MLOps practices and technologies such as CI/CD, IaC, and containerization.
    4. Model monitoring and maintenance: In this step, the deployed model is continuously monitored and maintained using MLOps practices and technologies such as model monitoring, automated testing, and container orchestration.

    MLOps is a rapidly growing field that has become essential for businesses that want to derive value from their machine learning investments. By leveraging MLOps practices and technologies, organizations can streamline the machine learning lifecycle, reduce deployment times, improve performance, and enhance security and compliance. While MLOps is still a relatively new field, it has the potential to transform the way that organizations build and deploy machine learning models, and to drive innovation and growth across industries.

    Link – https://medium.com/@Abhishek_Srivastava/mlops-bridging-the-gap-between-machine-learning-and-operations-abb47b5c03aa