AI in Practice Survey 2026
In running our second annual AI in Practice Survey, it is clear the stack has consolidated. Teams are wiring agents into their own systems on closed models, customizing models less, and checking quality with humans and vibes rather than rigorous evals. MCPs won as the standard for tool calling.
We surveyed 455 technical builders to map where adoption is happening, where the gaps are, and how teams are hardening and scaling AI in practice. Here’s what we found.
Key Findings
Security is the biggest pain point, with the fewest solutions.
Security rose from sixth to first among pain points between 2025 and 2026, and it has the widest gap between pain and solutions. Builders expect security tooling to deliver far more value than it does today, suggesting ample opportunity for new solutions.
43% rate security and governance 6 or 7 out of 7, double the next pain point (evals, 22%). Only 10% say security tooling delivers the most value today, while 18% expect it to deliver the most value in the future, the biggest expected gain of any area.
Security is the biggest pain point, with the fewest solutions.
43% rate security and governance 6 or 7 out of 7, double the next pain point (evals, 22%). Only 10% say security tooling delivers the most value today, while 18% expect it to deliver the most value in the future, the biggest expected gain of any area.
Where each area ranks for pain, 2025 vs 2026
Where tooling delivers value today vs where builders expect it in the future, 2026
Source: Q21, Q22, Q23 · 2025 n=413, 2026 n=455
2025 asked respondents to rank the seven areas; 2026 asked for a 1–7 rating of each. The top panel compares each area’s place by average, not the scores. User acceptance and Storage tied for 2nd in 2025.
Cost was the most common write-in pain point (9 of 30), but it wasn’t an answer option.
| Area | Rate pain 6–7 | Average pain (1–7) | Most value today | Most exciting (2025) | Most exciting (2026) | Pain rank (2025) | Pain rank (2026) |
|---|---|---|---|---|---|---|---|
| Security & governance | 43% | 4.9 | 10% | 10% | 18% | 6 | 1 |
| Evals | 22% | 4.5 | 20% | 25% | 21% | 7 | 2 |
| Data review | 18% | 4.3 | 20% | 21% | 20% | 5 | 3 |
| Context | 14% | 4.1 | 17% | 7% | 15% | 1 | 4 |
| User acceptance | 16% | 4.0 | 14% | 10% | 11% | 2 (tied) | 5 |
| Model management | 8% | 3.8 | 11% | 8% | 11% | 4 | 6 |
| Storage | 7% | 3.7 | 7% | 18% | 4% | 2 (tied) | 7 |
MCP became the standard for internal tools.
The year of the MCP! MCP won the standards race against A2A, though direct tool calling is slightly more common (61% vs 57%). And MCPs turned inward: teams building their own now mostly build them for internal users.
Teams using MCP to connect chat clients to tools and data rose from 33% to 57%, while A2A fell from 17% to 8%. Among teams building their own MCPs, the share building only for internal users rose from 44% to 67%, driven by non-tech teams (36% to 73%); tech teams barely moved (53% to 49%).
MCP became the standard for internal tools.
Teams using MCP to connect chat clients to tools and data rose from 33% to 57%, while A2A fell from 17% to 8%. Among teams building their own MCPs, the share building only for internal users rose from 44% to 67%, driven by non-tech teams (36% to 73%); tech teams barely moved (53% to 49%).
Who teams building their own MCPs build them for
MCP vs A2A: how teams connect chat clients to tools and data
Source: SQ4 (2025 Q1), Q17, Q20 · 2026 n=455 · 2025 n=413
Top: each team building its own MCPs is counted once, by who it builds them for. Teams naming only “Other” are excluded.
2026 has more non-tech MCP builders than 2025 (n=146 vs 69), and they drive the shift; tech teams barely moved.
Bottom: 2025 allowed one answer and 2026 several, which lifts every 2026 share. A2A fell anyway.
| Measure | 2025 | 2026 |
|---|---|---|
| Building MCP servers (share of all teams) | 17% | 42% |
| Connect chat clients via MCP (share of all teams) | 33% | 57% |
| Connect chat clients via A2A (share of all teams) | 17% | 8% |
| Connect chat clients via Direct tool calling (share of all teams) | 36% | 61% |
| Connect chat clients via No current usage (share of all teams) | 14% | 17% |
| Teams building MCPs: internal users only (n=135 → 203) | 44% | 67% |
| Teams building MCPs: internal and external (n=135 → 203) | 47% | 27% |
| Teams building MCPs: external users only (n=135 → 203) | 9% | 7% |
| Tech teams building MCPs: internal users only (n=66 → 57) | 53% | 49% |
| Tech teams building MCPs: internal and external (n=66 → 57) | 39% | 40% |
| Tech teams building MCPs: external users only (n=66 → 57) | 8% | 11% |
| Non-tech teams building MCPs: internal users only (n=69 → 146) | 36% | 73% |
| Non-tech teams building MCPs: internal and external (n=69 → 146) | 54% | 21% |
| Non-tech teams building MCPs: external users only (n=69 → 146) | 10% | 5% |
Open source lost the year.
Our 2025 survey surprised us with the preeminence of open source models. While there’s significant discourse about the need for enterprises to protect their alpha from frontier lab offerings, the data is telling a different story: ease of use and concerns about open source have led to a shift towards closed models.
Teams leaning mostly or only open source fell from 66% to 41%, while closed-leaning teams rose from 25% to 55%. Most teams still mix both (63% in 2025, 65% in 2026); what flipped is which side dominates. Tech teams are now the most closed (63%), and the shift holds in every company size and sector.
Open source lost the year.
Teams leaning mostly or only open source fell from 66% to 41%, while closed-leaning teams rose from 25% to 55%. Most teams still mix both (63% in 2025, 65% in 2026); what flipped is which side dominates. Tech teams are now the most closed (63%), and the shift holds in every company size and sector.
Source: Q10 · 2026 n=455 · 2025 n=413
| Model strategy | 2025 | 2026 |
|---|---|---|
| Only open source | 22% | 10% |
| Mostly open, some closed | 45% | 31% |
| Even split | 8% | 4% |
| Mostly closed, some open | 10% | 30% |
| Only closed source | 15% | 25% |
| Open-leaning (only + mostly open) | 66% | 41% |
| Closed-leaning (only + mostly closed), All teams | 25% | 55% |
| Closed-leaning (only + mostly closed), Tech teams | 25% | 63% |
| Closed-leaning (only + mostly closed), Other sectors | 26% | 53% |
| Closed-leaning (only + mostly closed), Under 1,000 employees | 24% | 54% |
| Closed-leaning (only + mostly closed), 1,000+ employees | 29% | 56% |
Half of AI teams have no eval showing they beat ChatGPT.
Only 49% use predefined evals to check whether their product beats using the frontier models directly. The rest rely on customer feedback, interviews and dogfooding, and 8% admit they’re not better at their target tasks.
Half of AI teams have no eval showing they beat ChatGPT.
Only 49% use predefined evals to check whether their product beats using the frontier models directly. The rest rely on customer feedback, interviews and dogfooding, and 8% admit they’re not better at their target tasks.
Source: Q2 (new in 2026) · n=455, each dot is one respondent
| Method | Teams | Share |
|---|---|---|
| Organic customer feedback | 258 | 57% |
| User-feedback interviews | 258 | 57% |
| Predefined evals | 222 | 49% |
| Dogfooding / smoke tests | 240 | 53% |
| Other | 19 | 4% |
| Not better than frontier models | 36 | 8% |
| No predefined evals (any other method) | 233 | 51% |
| Predefined evals, tech | 64% | |
| Predefined evals, other | 45% | |
| Predefined evals, 1-10 | 24% | |
| Predefined evals, 11-200 | 51% | |
| Predefined evals, 201-1000 | 45% | |
| Predefined evals, 1001+ | 52% |
Fine-tuning is fading, and what’s left is DIY.
Have models improved to an asymptote that precludes the need for fine-tuning? Respondents suggest this is the case, with overall fine-tuning utilization dropping, particularly with off-the-shelf products. Where fine-tuning occurs, DIY is most common.
Teams that have ever fine-tuned fell from 81% to 51%. OpenAI’s fine-tuning product fell from 54% to 17%, while homegrown libraries rose from 7% to 20% and now make up 40% of all fine-tuners. The drop is steepest at companies with 1,000 or fewer employees (82% to 38%).
Fine-tuning is fading, and what’s left is DIY.
Teams that have ever fine-tuned fell from 81% to 51%. OpenAI’s fine-tuning product fell from 54% to 17%, while homegrown libraries rose from 7% to 20% and now make up 40% of all fine-tuners. The drop is steepest at companies with 1,000 or fewer employees (82% to 38%).
Source: Q11 · 2026 n=455 · 2025 n=413 · each square is 1% of all teams
| Group | Category | 2025 | 2026 |
|---|---|---|---|
| All | Ever fine-tuned | 81% | 51% |
| All | OpenAI RFT product | 54% | 17% |
| All | Other paid service | 20% | 14% |
| All | Homegrown library | 7% | 20% |
| All | Never fine-tuned | 19% | 49% |
| Under 1,000 | Ever fine-tuned | 82% | 38% |
| Under 1,000 | OpenAI RFT product | 57% | 9% |
| Under 1,000 | Other paid service | 20% | 10% |
| Under 1,000 | Homegrown library | 5% | 19% |
| Under 1,000 | Never fine-tuned | 18% | 62% |
| 1,000+ | Ever fine-tuned | 79% | 61% |
| 1,000+ | OpenAI RFT product | 48% | 23% |
| 1,000+ | Other paid service | 20% | 17% |
| 1,000+ | Homegrown library | 11% | 21% |
| 1,000+ | Never fine-tuned | 21% | 39% |
| 1–10 employees | Ever fine-tuned | 48% | 24% |
| 1–10 employees | OpenAI RFT product | 29% | 0% |
| 1–10 employees | Other paid service | 14% | 5% |
| 1–10 employees | Homegrown library | 5% | 19% |
| 1–10 employees | Never fine-tuned | 52% | 76% |
| 11–200 employees | Ever fine-tuned | 83% | 39% |
| 11–200 employees | OpenAI RFT product | 58% | 7% |
| 11–200 employees | Other paid service | 18% | 12% |
| 11–200 employees | Homegrown library | 7% | 19% |
| 11–200 employees | Never fine-tuned | 17% | 61% |
| 201–1,000 employees | Ever fine-tuned | 86% | 40% |
| 201–1,000 employees | OpenAI RFT product | 60% | 12% |
| 201–1,000 employees | Other paid service | 22% | 9% |
| 201–1,000 employees | Homegrown library | 4% | 19% |
| 201–1,000 employees | Never fine-tuned | 14% | 60% |
Coding agents are partners in observability.
Given the scale and difficulty of making sense of log files, developers are more than happy to delegate this task to coding agents: 48% of teams now point coding agents at their production logs. That said, few use coding agents alone: only 41 of the 219 coding-agent teams use them with no observability, BI or spreadsheet tool. Meanwhile, tech teams continue to use observability tools (53%), and tech teams storing full traces there rose from 57% to 64%. The overall decline of trace logging, from 81% to 73%, mostly comes from non-tech and smaller companies.
The tool that lost out is the spreadsheet (57% → 28%). Teams using coding agents are more likely to also use LLM observability tools (47% vs 35%), not less.
Coding agents are partners in observability.
The tool that lost out is the spreadsheet (57% → 28%). Teams using coding agents are more likely to also use LLM observability tools (47% vs 35%), not less.
Source: Q5, Q4 · 2026 n=455 · 2025 n=413
The coding-agent option was added in 2026; part of the spreadsheet decline may reflect people choosing the new option rather than a true replacement.
| Measure | 2025 | 2026 |
|---|---|---|
| Review: Spreadsheets | 57% | 28% |
| Review: BI / data science | 51% | 56% |
| Review: LLM observability | 53% | 41% |
| Review: Don’t review | 2% | 7% |
| Review: coding agents analyzing logs (new in 2026) | Not asked | 48% |
| Traces: Full traces in observability product | 56% | 45% |
| Traces: Any explicit trace logging | 81% | 73% |
| Traces: Full traces, tech only | 57% | 64% |
| Traces: Full traces, other sectors | 56% | 41% |
| Traces: Any trace logging, tech only | 80% | 85% |
| Traces: Any trace logging, other sectors | 81% | 70% |
Humans and vibes, not LLM judges, run evals.
Companies took evals in-house, while labeling vendors’ demand moved to frontier labs. Could cost be driving down data-labeling vendors for evaluating LLMs? We see a major shift in data-labeling usage (17% to 2%) while human feedback is up.
Human feedback is used by 86%, automated evals by 47% and LLM-as-judge by 38%. Internal expert reviewers rose from 27% to 54%, data-labeling vendors collapsed from 17% to 2%, and “vibes” as a quality check rose from 41% to 48%.
Humans and vibes, not LLM judges, run evals.
Human feedback is used by 86%, automated evals by 47% and LLM-as-judge by 38%. Internal expert reviewers rose from 27% to 54%, data-labeling vendors collapsed from 17% to 2%, and “vibes” as a quality check rose from 41% to 48%.
Source: Q7, Q8 (methods, human review), Q6 (quality checks) · 2026 n=455 · 2025 n=413
“Who does the human review?” is a share of teams using human feedback; in 2025 that question went to everyone, so it is re-based to teams that selected human feedback.
| Measure | Base | 2025 | 2026 |
|---|---|---|---|
| Human feedback | of all teams | 74% | 86% |
| Automated evals | of all teams | 63% | 47% |
| LLM-as-judge | of all teams | 40% | 38% |
| Internal experts | of teams using human feedback (n=303 → 390) | 27% | 54% |
| Internal + external mix | of teams using human feedback (n=303 → 390) | 56% | 43% |
| Data-labeling companies | of teams using human feedback (n=303 → 390) | 17% | 2% |
| Vibes (subjective judgment) | of all teams | 41% | 48% |
| Telemetry | of all teams | 24% | 32% |
| AI evaluations | of all teams | 67% | 60% |
Big companies run the most advanced AI stacks, but not yet the best returns.
Enterprises are driving utilization across AI techniques, demonstrating a maturation of the adoption and product availability for these customers. Companies with 1,000+ employees lead smaller ones on code mode, routing, skills, and fine-tuning. That said, we’re hearing that many companies are not yet showing ROI on these investments.
Big companies run the most advanced AI stacks, but not yet the best returns.
Enterprises are driving utilization across AI techniques, demonstrating a maturation of the adoption and product availability for these customers. Companies with 1,000+ employees lead smaller ones on code mode, routing, skills, and fine-tuning. That said, we’re hearing that many companies are not yet showing ROI on these investments.
Source: Q11, Q12, Q15, Q16 by company size (D4) · 2026 n=455 (1,000+ n=248, under 1,000 n=207)
Compares company sizes within 2026, not change over time. “Startups” means companies under 1,000 employees.
| Technique | Under 1,000 | 1,000+ | Gap | 1–10 | 11–200 | 201–1,000 | 1,000+ |
|---|---|---|---|---|---|---|---|
| Code mode | 50% | 63% | +13 pts | 29% (6 of 21) | 55% (37 of 67) | 51% (61 of 119) | 63% (155 of 248) |
| Routing | 35% | 54% | +19 pts | 24% (5 of 21) | 39% (26 of 67) | 35% (42 of 119) | 54% (133 of 248) |
| Skills | 74% | 84% | +10 pts | 67% (14 of 21) | 79% (53 of 67) | 72% (86 of 119) | 84% (208 of 248) |
| Fine-tuning | 38% | 61% | +23 pts | 24% (5 of 21) | 39% (26 of 67) | 40% (48 of 119) | 61% (152 of 248) |
Survey Results
Full results for all 2026 questions. Filter to compare segments; shares are of respondents who were asked each question.
Respondents
455 responses
| Response | Share |
|---|---|
| I manage or lead teams that build / deploy AI tools / applications | 86.8% |
| I help deploy, test, or integrate AI tools / applications | 63.3% |
| I directly build or develop AI tools / applications | 54.7% |
| I only use AI tools built by others | 8.4% |
455 responses
| Response | Share |
|---|---|
| 1-10 | 4.6% |
| 11-200 | 14.7% |
| 201-1000 | 26.2% |
| 1001+ | 54.5% |
455 responses
| Response | Share |
|---|---|
| Technology and Software | 19.6% |
| Government and Public sector | 18.7% |
| Healthcare and Life Sciences | 17.8% |
| Financial services and Insurance | 15.8% |
| Manufacturing, Energy & Industrial | 14.9% |
| Media, Telecom & Entertainment | 6.8% |
| Retail, eCommerce & Consumer Goods | 6.4% |
455 responses
| Response | Share |
|---|---|
| Business owner | 3.1% |
| Business partner / co-owner | 4.2% |
| C-suite (not the owner) | 40.7% |
| Senior level management | 52.1% |
What teams are building
455 responses
| Response | Share |
|---|---|
| Internal user-facing AI features (incl. tools and workflows) | 90.3% |
| AI that improves existing products | 73.4% |
| External user-facing AI features | 63.1% |
| Systems for enabling AI products | 51.6% |
| MCP servers | 41.5% |
| Other | 1.5% |
455 responses
| Response | Share |
|---|---|
| Organic customer feedback and product signals | 56.7% |
| User-feedback interviews | 56.7% |
| Ad-hoc evaluation by internal dogfooding or smoke tests | 52.7% |
| Predefined evals | 48.8% |
| We currently are not better than the frontier models at our target tasks | 7.9% |
| Other | 4.2% |
Models & customization
455 responses
| Response | Share |
|---|---|
| Only open source | 10.1% |
| Mostly open source | 31.0% |
| Even split | 4.2% |
| Mostly closed source | 30.1% |
| Only closed source | 24.6% |
455 responses
| Response | Share |
|---|---|
| Yes, via OpenAI RFT | 16.7% |
| Yes, via another paid service | 13.6% |
| Yes, via a homegrown library | 20.4% |
| No | 49.2% |
455 responses
| Response | Share |
|---|---|
| Yes, we have a custom router we trained / built | 23.7% |
| Yes, we’re using a proprietary router | 21.5% |
| No, all user queries use the same model | 54.7% |
Agents & code
455 responses
| Response | Share |
|---|---|
| Database access | 78.9% |
| First party API / MCP servers | 67.0% |
| Web search | 64.6% |
| File systems | 62.6% |
| Plotting / dashboarding tools | 46.2% |
| Code interpreter | 43.7% |
| Memory systems | 40.0% |
| Other | 1.8% |
455 responses
| Response | Share |
|---|---|
| Yes — autonomous coding agents in a sandboxed harness | 23.7% |
| Yes — tool calls to sandboxed coding subagents | 33.2% |
| No — our agents can’t write and execute code | 31.0% |
| We don’t use autonomous agents yet | 12.1% |
259 responses
| Response | Share |
|---|---|
| Cloudflare | 62.2% |
| Other | 17.4% |
| Vercel | 16.6% |
| Modal | 15.8% |
| E2B | 12.4% |
| Daytona | 8.5% |
| Sail | 6.6% |
Asked of teams whose agents write and execute code.
259 responses
| Response | Share |
|---|---|
| Claude | 76.8% |
| ChatGPT | 39.0% |
| Codex | 33.2% |
| Gemini | 29.3% |
| Cursor | 23.9% |
| Opencode | 12.0% |
| Pi | 3.9% |
| Other | 3.9% |
Asked of teams whose agents write and execute code.
Data & observability
455 responses
| Response | Share |
|---|---|
| We are using a database and logging traces as rows | 50.5% |
| We are storing everything as traces in a product designed for LLM observability | 45.3% |
| We are storing user prompts and feedback but not the full traces | 40.0% |
| We are storing the responses of tool calls | 35.6% |
| We are not storing this data | 7.0% |
| Other | 2.2% |
455 responses
| Response | Share |
|---|---|
| We use a traditional Business Intelligence / Data Science tool | 55.8% |
| We pull the logs and analyze them with coding agents | 48.1% |
| We use a specialized LLM observability tool | 41.1% |
| We use spreadsheets to look at the data and extract insights | 28.4% |
| We do not currently review the data as part of our process | 7.0% |
| Other | 1.5% |
Quality & evaluation
455 responses
| Response | Share |
|---|---|
| AI evaluations (manual or automated) | 60.2% |
| User satisfaction surveys | 54.7% |
| A/B testing with users | 48.4% |
| Subjective qualitative judgment (i.e., “Vibes”) | 48.1% |
| Telemetry | 32.3% |
| We aren’t yet looking at quality | 9.9% |
| Other | 2.0% |
455 responses
| Response | Share |
|---|---|
| Using human feedback | 85.7% |
| Using automated evals | 47.3% |
| Using LLM-as-a-judge | 37.6% |
| N/A - we are not performing AI evaluation | 3.7% |
| Other | 1.5% |
| None of the above | 0.7% |
390 responses
| Response | Share |
|---|---|
| Have internal experts reviewing the results | 54.4% |
| A combination of internal and external reviewers | 42.6% |
| Working with data labeling companies | 2.3% |
| Other | 0.8% |
Asked of teams using human feedback.
215 responses
| Response | Share |
|---|---|
| Doing error analysis on the failure cases | 64.2% |
| Running them as regression tests | 62.3% |
| Using verification strategies | 60.9% |
| Running them as unit tests | 56.7% |
| Other | 0.5% |
Asked of teams using automated evals.
MCP & skills
455 responses
| Response | Share |
|---|---|
| API via tools | 60.7% |
| MCP | 57.4% |
| No current usage | 16.7% |
| A2A | 8.1% |
261 responses
| Response | Share |
|---|---|
| Building internal MCPs to access data in your own systems | 69.7% |
| Using first-party MCPs published by other software providers (e.g. Notion, MotherDuck) | 53.6% |
| Building internal MCPs to access external APIs | 34.5% |
| Other | 0.8% |
Asked of teams using MCP.
205 responses
| Response | Share |
|---|---|
| Primarily using them to provide an API interface for a chat client | 43.9% |
| Building custom logic and context layers for APIs | 42.4% |
| Integrating with advanced MCP features like Widgets / UI | 13.2% |
| Other | 0.5% |
Asked of teams building their own MCPs.
205 responses
| Response | Share |
|---|---|
| Internal technical users (e.g. engineers) | 76.6% |
| Internal less-technical users (e.g. business analysts) | 64.9% |
| External users (e.g. consumers) | 33.2% |
| Other | 1.5% |
Asked of teams building their own MCPs.
455 responses
| Response | Share |
|---|---|
| Yes, we expose skills as part of an MCP | 41.8% |
| Yes, we expose skills as local files deployed with the agent | 36.0% |
| Yes, we expose skills via another mechanism | 1.5% |
| No | 20.7% |
Pain points & value
455 responses · bar = average rating · % = share rating 6 or 7
| Area | Average rating | Rated 6 or 7 |
|---|---|---|
| Security & governance | 4.9 | 43.1% |
| Data quality assessment (evals) | 4.5 | 22.0% |
| Data review of interactions with our production AI | 4.3 | 17.8% |
| Context | 4.1 | 14.1% |
| User acceptance | 4.0 | 16.3% |
| Model management | 3.8 | 7.7% |
| Storage of interactions with our AI | 3.7 | 6.8% |
455 responses
| Response | Share |
|---|---|
| Data review of interactions with our production AI | 20.4% |
| Data Quality Assessment (Evals) | 20.0% |
| Context | 17.4% |
| User acceptance | 13.6% |
| Model management | 10.5% |
| Security & governance | 10.3% |
| Storage of interactions with our AI | 7.3% |
| Other | 0.4% |
455 responses
| Response | Share |
|---|---|
| Data Quality Assessment (Evals) | 20.7% |
| Data review of interactions with our production AI | 19.8% |
| Security & governance | 17.6% |
| Context | 15.4% |
| Model management | 11.2% |
| User acceptance | 10.5% |
| Storage of interactions with our AI | 4.4% |
| Other | 0.4% |
About this survey
455 respondents, fielded September 22, 2026. Percentages are rounded; multi-select questions sum to more than 100%.
Notes on respondent mix: All respondents build, deploy, or lead teams that build AI applications at their organization. Our respondents are more senior and more enterprise in 2026 than our 2025 survey (n = 413): 55% work at companies with 1,001+ employees and 41% are C-suite. Every year-over-year shift highlighted in the findings also holds, in direction and rough magnitude, when both years are restricted to companies with 201+ employees.