AI in Practice Survey 2026

In running our second annual AI in Practice Survey, it is clear the stack has consolidated. Teams are wiring agents into their own systems on closed models, customizing models less, and checking quality with humans and vibes rather than rigorous evals. MCPs won as the standard for tool calling.

We surveyed 455 technical builders to map where adoption is happening, where the gaps are, and how teams are hardening and scaling AI in practice. Here’s what we found.

Key Findings

1

Security is the biggest pain point, with the fewest solutions.

Security rose from sixth to first among pain points between 2025 and 2026, and it has the widest gap between pain and solutions. Builders expect security tooling to deliver far more value than it does today, suggesting ample opportunity for new solutions.

43% rate security and governance 6 or 7 out of 7, double the next pain point (evals, 22%). Only 10% say security tooling delivers the most value today, while 18% expect it to deliver the most value in the future, the biggest expected gain of any area.

View supporting data →
2

MCP became the standard for internal tools.

The year of the MCP! MCP won the standards race against A2A, though direct tool calling is slightly more common (61% vs 57%). And MCPs turned inward: teams building their own now mostly build them for internal users.

Teams using MCP to connect chat clients to tools and data rose from 33% to 57%, while A2A fell from 17% to 8%. Among teams building their own MCPs, the share building only for internal users rose from 44% to 67%, driven by non-tech teams (36% to 73%); tech teams barely moved (53% to 49%).

View supporting data →
3

Open source lost the year.

Our 2025 survey surprised us with the preeminence of open source models. While there’s significant discourse about the need for enterprises to protect their alpha from frontier lab offerings, the data is telling a different story: ease of use and concerns about open source have led to a shift towards closed models.

Teams leaning mostly or only open source fell from 66% to 41%, while closed-leaning teams rose from 25% to 55%. Most teams still mix both (63% in 2025, 65% in 2026); what flipped is which side dominates. Tech teams are now the most closed (63%), and the shift holds in every company size and sector.

View supporting data →
4

Half of AI teams have no eval showing they beat ChatGPT.

Only 49% use predefined evals to check whether their product beats using the frontier models directly. The rest rely on customer feedback, interviews and dogfooding, and 8% admit they’re not better at their target tasks.

View supporting data →
5

Fine-tuning is fading, and what’s left is DIY.

Have models improved to an asymptote that precludes the need for fine-tuning? Respondents suggest this is the case, with overall fine-tuning utilization dropping, particularly with off-the-shelf products. Where fine-tuning occurs, DIY is most common.

Teams that have ever fine-tuned fell from 81% to 51%. OpenAI’s fine-tuning product fell from 54% to 17%, while homegrown libraries rose from 7% to 20% and now make up 40% of all fine-tuners. The drop is steepest at companies with 1,000 or fewer employees (82% to 38%).

View supporting data →
6

Coding agents are partners in observability.

Given the scale and difficulty of making sense of log files, developers are more than happy to delegate this task to coding agents: 48% of teams now point coding agents at their production logs. That said, few use coding agents alone: only 41 of the 219 coding-agent teams use them with no observability, BI or spreadsheet tool. Meanwhile, tech teams continue to use observability tools (53%), and tech teams storing full traces there rose from 57% to 64%. The overall decline of trace logging, from 81% to 73%, mostly comes from non-tech and smaller companies.

The tool that lost out is the spreadsheet (57% → 28%). Teams using coding agents are more likely to also use LLM observability tools (47% vs 35%), not less.

View supporting data →
7

Humans and vibes, not LLM judges, run evals.

Companies took evals in-house, while labeling vendors’ demand moved to frontier labs. Could cost be driving down data-labeling vendors for evaluating LLMs? We see a major shift in data-labeling usage (17% to 2%) while human feedback is up.

Human feedback is used by 86%, automated evals by 47% and LLM-as-judge by 38%. Internal expert reviewers rose from 27% to 54%, data-labeling vendors collapsed from 17% to 2%, and “vibes” as a quality check rose from 41% to 48%.

View supporting data →
8

Big companies run the most advanced AI stacks, but not yet the best returns.

Enterprises are driving utilization across AI techniques, demonstrating a maturation of the adoption and product availability for these customers. Companies with 1,000+ employees lead smaller ones on code mode, routing, skills, and fine-tuning. That said, we’re hearing that many companies are not yet showing ROI on these investments.

View supporting data →
Appendix

Survey Results

Full results for all 2026 questions. Filter to compare segments; shares are of respondents who were asked each question.

Filter Responses

455 of 455

Respondents

Q1
Which best describes your involvement with building and deploying AI tools and applications?

455 responses

Which best describes your involvement with building and deploying AI tools and applications? (455 responses)
ResponseShare
I manage or lead teams that build / deploy AI tools / applications86.8%
I help deploy, test, or integrate AI tools / applications63.3%
I directly build or develop AI tools / applications54.7%
I only use AI tools built by others8.4%
Q2
Approximately how many employees are at your company?

455 responses

Approximately how many employees are at your company? (455 responses)
ResponseShare
1-104.6%
11-20014.7%
201-100026.2%
1001+54.5%
Q3
What sector is your company in?

455 responses

What sector is your company in? (455 responses)
ResponseShare
Technology and Software19.6%
Government and Public sector18.7%
Healthcare and Life Sciences17.8%
Financial services and Insurance15.8%
Manufacturing, Energy & Industrial14.9%
Media, Telecom & Entertainment6.8%
Retail, eCommerce & Consumer Goods6.4%
Q4
Which best describes your current level of employment?

455 responses

Which best describes your current level of employment? (455 responses)
ResponseShare
Business owner3.1%
Business partner / co-owner4.2%
C-suite (not the owner)40.7%
Senior level management52.1%

What teams are building

Q5
What kinds of AI products are you building?

455 responses

What kinds of AI products are you building? (455 responses)
ResponseShare
Internal user-facing AI features (incl. tools and workflows)90.3%
AI that improves existing products73.4%
External user-facing AI features63.1%
Systems for enabling AI products51.6%
MCP servers41.5%
Other1.5%
Q6New in 2026
How do you determine if your product is better than directly using the frontier models (ChatGPT, Claude, Gemini)?

455 responses

How do you determine if your product is better than directly using the frontier models (ChatGPT, Claude, Gemini)? (455 responses)
ResponseShare
Organic customer feedback and product signals56.7%
User-feedback interviews56.7%
Ad-hoc evaluation by internal dogfooding or smoke tests52.7%
Predefined evals48.8%
We currently are not better than the frontier models at our target tasks7.9%
Other4.2%

Models & customization

Q7
Are you using open source or closed source models for the AI capabilities your team is building?

455 responses

Are you using open source or closed source models for the AI capabilities your team is building? (455 responses)
ResponseShare
Only open source10.1%
Mostly open source31.0%
Even split4.2%
Mostly closed source30.1%
Only closed source24.6%
Q8
Have you tried fine-tuning?

455 responses

Have you tried fine-tuning? (455 responses)
ResponseShare
Yes, via OpenAI RFT16.7%
Yes, via another paid service13.6%
Yes, via a homegrown library20.4%
No49.2%
Q9New in 2026
Are you doing any routing?

455 responses

Are you doing any routing? (455 responses)
ResponseShare
Yes, we have a custom router we trained / built23.7%
Yes, we’re using a proprietary router21.5%
No, all user queries use the same model54.7%

Agents & code

Q10
What kind of tools are your agents using?

455 responses

What kind of tools are your agents using? (455 responses)
ResponseShare
Database access78.9%
First party API / MCP servers67.0%
Web search64.6%
File systems62.6%
Plotting / dashboarding tools46.2%
Code interpreter43.7%
Memory systems40.0%
Other1.8%
Q11New in 2026
Does your AI application use a coding agent in a harness (“code mode”)?

455 responses

Does your AI application use a coding agent in a harness (“code mode”)? (455 responses)
ResponseShare
Yes — autonomous coding agents in a sandboxed harness23.7%
Yes — tool calls to sandboxed coding subagents33.2%
No — our agents can’t write and execute code31.0%
We don’t use autonomous agents yet12.1%
Q12New in 2026
Where are the sandboxes hosted?

259 responses

Where are the sandboxes hosted? (259 responses)
ResponseShare
Cloudflare62.2%
Other17.4%
Vercel16.6%
Modal15.8%
E2B12.4%
Daytona8.5%
Sail6.6%

Asked of teams whose agents write and execute code.

Q13New in 2026
Which coding harness are you using?

259 responses

Which coding harness are you using? (259 responses)
ResponseShare
Claude76.8%
ChatGPT39.0%
Codex33.2%
Gemini29.3%
Cursor23.9%
Opencode12.0%
Pi3.9%
Other3.9%

Asked of teams whose agents write and execute code.

Data & observability

Q14
How are you storing interactions with your AI system?

455 responses

How are you storing interactions with your AI system? (455 responses)
ResponseShare
We are using a database and logging traces as rows50.5%
We are storing everything as traces in a product designed for LLM observability45.3%
We are storing user prompts and feedback but not the full traces40.0%
We are storing the responses of tool calls35.6%
We are not storing this data7.0%
Other2.2%
Q15
How are you reviewing the data of interactions with your system?

455 responses

How are you reviewing the data of interactions with your system? (455 responses)
ResponseShare
We use a traditional Business Intelligence / Data Science tool55.8%
We pull the logs and analyze them with coding agents48.1%
We use a specialized LLM observability tool41.1%
We use spreadsheets to look at the data and extract insights28.4%
We do not currently review the data as part of our process7.0%
Other1.5%

Quality & evaluation

Q16
How are you assessing the quality of your AI products?

455 responses

How are you assessing the quality of your AI products? (455 responses)
ResponseShare
AI evaluations (manual or automated)60.2%
User satisfaction surveys54.7%
A/B testing with users48.4%
Subjective qualitative judgment (i.e., “Vibes”)48.1%
Telemetry32.3%
We aren’t yet looking at quality9.9%
Other2.0%
Q17
If you’re doing AI evaluations, which of the following are you using?

455 responses

If you’re doing AI evaluations, which of the following are you using? (455 responses)
ResponseShare
Using human feedback85.7%
Using automated evals47.3%
Using LLM-as-a-judge37.6%
N/A - we are not performing AI evaluation3.7%
Other1.5%
None of the above0.7%
Q18
Which of the following are you using with human feedback?

390 responses

Which of the following are you using with human feedback? (390 responses)
ResponseShare
Have internal experts reviewing the results54.4%
A combination of internal and external reviewers42.6%
Working with data labeling companies2.3%
Other0.8%

Asked of teams using human feedback.

Q19
Which of the following are you using with automated evaluations?

215 responses

Which of the following are you using with automated evaluations? (215 responses)
ResponseShare
Doing error analysis on the failure cases64.2%
Running them as regression tests62.3%
Using verification strategies60.9%
Running them as unit tests56.7%
Other0.5%

Asked of teams using automated evals.

MCP & skills

Q20
Are you using LLM chat clients to access APIs or data sources via MCP or direct tool calling?

455 responses

Are you using LLM chat clients to access APIs or data sources via MCP or direct tool calling? (455 responses)
ResponseShare
API via tools60.7%
MCP57.4%
No current usage16.7%
A2A8.1%
Q21
How are you using MCPs?

261 responses

How are you using MCPs? (261 responses)
ResponseShare
Building internal MCPs to access data in your own systems69.7%
Using first-party MCPs published by other software providers (e.g. Notion, MotherDuck)53.6%
Building internal MCPs to access external APIs34.5%
Other0.8%

Asked of teams using MCP.

Q22
How are you building your own MCPs?

205 responses

How are you building your own MCPs? (205 responses)
ResponseShare
Primarily using them to provide an API interface for a chat client43.9%
Building custom logic and context layers for APIs42.4%
Integrating with advanced MCP features like Widgets / UI13.2%
Other0.5%

Asked of teams building their own MCPs.

Q23
Who is your intended audience for the MCPs you are building?

205 responses

Who is your intended audience for the MCPs you are building? (205 responses)
ResponseShare
Internal technical users (e.g. engineers)76.6%
Internal less-technical users (e.g. business analysts)64.9%
External users (e.g. consumers)33.2%
Other1.5%

Asked of teams building their own MCPs.

Q24New in 2026
Are you utilizing skills for your LLM application?

455 responses

Are you utilizing skills for your LLM application? (455 responses)
ResponseShare
Yes, we expose skills as part of an MCP41.8%
Yes, we expose skills as local files deployed with the agent36.0%
Yes, we expose skills via another mechanism1.5%
No20.7%

Pain points & value

Q25
Where are you experiencing the biggest pain points today in deploying your AI products?

455 responses · bar = average rating · % = share rating 6 or 7

Where are you experiencing the biggest pain points today in deploying your AI products?
AreaAverage ratingRated 6 or 7
Security & governance4.943.1%
Data quality assessment (evals)4.522.0%
Data review of interactions with our production AI4.317.8%
Context4.114.1%
User acceptance4.016.3%
Model management3.87.7%
Storage of interactions with our AI3.76.8%
Q26
Which area of tooling is delivering the most value today for your projects?

455 responses

Which area of tooling is delivering the most value today for your projects? (455 responses)
ResponseShare
Data review of interactions with our production AI20.4%
Data Quality Assessment (Evals)20.0%
Context17.4%
User acceptance13.6%
Model management10.5%
Security & governance10.3%
Storage of interactions with our AI7.3%
Other0.4%
Q27
Which area of tooling are you most excited about for the value it may create in the future?

455 responses

Which area of tooling are you most excited about for the value it may create in the future? (455 responses)
ResponseShare
Data Quality Assessment (Evals)20.7%
Data review of interactions with our production AI19.8%
Security & governance17.6%
Context15.4%
Model management11.2%
User acceptance10.5%
Storage of interactions with our AI4.4%
Other0.4%
Methodology

About this survey

455 respondents, fielded September 22, 2026. Percentages are rounded; multi-select questions sum to more than 100%.

Notes on respondent mix: All respondents build, deploy, or lead teams that build AI applications at their organization. Our respondents are more senior and more enterprise in 2026 than our 2025 survey (n = 413): 55% work at companies with 1,001+ employees and 41% are C-suite. Every year-over-year shift highlighted in the findings also holds, in direction and rough magnitude, when both years are restricted to companies with 201+ employees.