In running our second annual AI in Practice Survey, it is clear the stack has consolidated. Teams are wiring agents into their own systems on closed models, customizing models less, and checking quality with humans and vibes rather than rigorous evals. MCPs won as the standard for tool calling.
We surveyed 455 technical builders to map where adoption is happening, where the gaps are, and how teams are hardening and scaling AI in practice. Here’s what we found.
Security rose from sixth to first among pain points between 2025 and 2026, and it has the widest gap between pain and solutions. Builders expect security tooling to deliver far more value than it does today, suggesting ample opportunity for new solutions.
43% rate security and governance 6 or 7 out of 7, double the next pain point (evals, 22%). Only 10% say security tooling delivers the most value today, while 18% expect it to deliver the most value in the future, the biggest expected gain of any area.

The year of the MCP! MCP won the standards race against A2A, though direct tool calling is slightly more common (61% vs 57%). And MCPs turned inward: teams building their own now mostly build them for internal users.
Teams using MCP to connect chat clients to tools and data rose from 33% to 57%, while A2A fell from 17% to 8%. Among teams building their own MCPs, the share building only for internal users rose from 44% to 67%, driven by non-tech teams (36% to 73%); tech teams barely moved (53% to 49%).

Our 2025 survey surprised us with the preeminence of open source models. While there’s significant discourse about the need for enterprises to protect their alpha from frontier lab offerings, the data is telling a different story: ease of use and concerns about open source have led to a shift towards closed models.
Teams leaning mostly or only open source fell from 66% to 41%, while closed-leaning teams rose from 25% to 55%. Most teams still mix both (63% in 2025, 65% in 2026); what flipped is which side dominates. Tech teams are now the most closed (63%), and the shift holds in every company size and sector.

Only 49% use predefined evals to check whether their product beats using the frontier models directly. The rest rely on customer feedback, interviews and dogfooding, and 8% admit they’re not better at their target tasks.

Have models improved to an asymptote that precludes the need for fine-tuning? Respondents suggest this is the case, with overall fine-tuning utilization dropping, particularly with off-the-shelf products. Where fine-tuning occurs, DIY is most common.
Teams that have ever fine-tuned fell from 81% to 51%. OpenAI’s fine-tuning product fell from 54% to 17%, while homegrown libraries rose from 7% to 20% and now make up 40% of all fine-tuners. The drop is steepest at companies with 1,000 or fewer employees (82% to 38%).

Given the scale and difficulty of making sense of log files, developers are more than happy to delegate this task to coding agents: 48% of teams now point coding agents at their production logs. That said, few use coding agents alone: only 41 of the 219 coding-agent teams use them with no observability, BI or spreadsheet tool. Meanwhile, tech teams continue to use observability tools (53%), and tech teams storing full traces there rose from 57% to 64%. The overall decline of trace logging, from 81% to 73%, mostly comes from non-tech and smaller companies.
The tool that lost out is the spreadsheet (57% → 28%). Teams using coding agents are more likely to also use LLM observability tools (47% vs 35%), not less.

Companies took evals in-house, while labeling vendors’ demand moved to frontier labs. Could cost be driving down data-labeling vendors for evaluating LLMs? We see a major shift in data-labeling usage (17% to 2%) while human feedback is up.
Human feedback is used by 86%, automated evals by 47% and LLM-as-judge by 38%. Internal expert reviewers rose from 27% to 54%, data-labeling vendors collapsed from 17% to 2%, and “vibes” as a quality check rose from 41% to 48%.

Enterprises are driving utilization across AI techniques, demonstrating a maturation of the adoption and product availability for these customers. Companies with 1,000+ employees lead smaller ones on code mode, routing, skills, and fine-tuning. That said, we’re hearing that many companies are not yet showing ROI on these investments.

Full results, all 27 questions, every chart: AI in Practice Survey 2026