On this page7 sections
MCP context window: the short answer
The MCP context window fills up because a client that loads tools up front sends every tool's name, description and input schema to the model before you type a word. In Anthropic's published example, 58 tools across five servers take about 55,000 tokens. Load fewer tools, or load them on demand with tool search.
Where the tokens go
A model has a fixed context window. Everything in a chat shares it:
- the system prompt and your instructions
- tool definitions from every connected MCP server
- your messages and the model's replies
- tool results, such as a list of 200 orders
Tool definitions are the part you never see. Unless the client defers them, they go out with every request, used or not.
How big the problem is
Anthropic published figures in its advanced tool use post of 24 November 2025:
| Server | Tools | Tokens (Anthropic's figures) |
|---|---|---|
| GitHub | 35 | ~26,000 |
| Slack | 11 | ~21,000 |
| Sentry | 5 | ~3,000 |
| Grafana | 5 | ~3,000 |
| Splunk | 2 | ~2,000 |
| Total | 58 | ~55,000 |
The same post puts Jira alone at about 17,000 tokens. It also says Anthropic has seen tool definitions take 134,000 tokens before optimization.
Those five servers are small next to a full business API. In PopMCP's catalog, HubSpot has 1,000+ operations, Stripe has 581 and Xero has 364. Loading a catalog that size up front is not practical, which is why large catalogs sit behind search.
Why too many MCP tools hurt
Less room for work. Every token spent on definitions is one less for documents, data and chat history.
Lower accuracy. The model has to pick one tool from a long list of similar names. Anthropic's tool search docs say selection gets worse past 30 to 50 available tools. In Anthropic's internal testing on large tool libraries, turning on tool search raised accuracy from 49% to 74% on Claude Opus 4, and from 79.5% to 88.1% on Opus 4.5.
Higher cost. Definitions count as input tokens. If you pay per token, you pay for them on each request.
Tool results add to the load. Claude Code warns when an MCP tool's output passes 10,000 tokens and caps it at 25,000 by default. One output at that cap is about the size of the GitHub server's 35 definitions in the table above.
MCP tool limits in popular clients
Checked in each vendor's own docs on 5 October 2026.
| Client | What its docs say |
|---|---|
| VS Code (GitHub Copilot) | A chat request can have at most 128 tools enabled. Past that you get the error "Cannot have more than 128 tools per request". |
| Claude Code | No fixed per-server tool cap. MCP tool search is on by default, so only tool names and server instructions load at session start. |
| Claude API | The tool search tool with defer_loading. Up to 10,000 deferred tools per request. |
| Cursor | The MCP docs state no tool cap. Since January 2026, Cursor syncs tool descriptions to a folder and gives the agent only tool names up front. |
Cursor calls its approach dynamic context discovery. In its A/B test, runs that called an MCP tool used 46.9% fewer total agent tokens. Cursor notes the result varied a lot with the number of servers installed.
These limits change, so check the client's docs before you plan a large setup.
How to fix MCP token usage
1. Connect fewer servers
Turn off the servers a task does not need. In Cursor, each server has a toggle under Customize. In VS Code, the tools picker lets you deselect single tools or whole servers. In Claude, the "Search and tools" menu lets you disable tools for the current conversation.
2. Use tool search
With tool search, the model gets one small search tool in place of every definition. It searches, and only the matching tools load. In Anthropic's example, total context use before any work fell from about 77,000 tokens to about 8,700. Anthropic reports that as an 85% reduction.
Where you get it:
- Claude API: add the tool search tool and set
defer_loading: trueper tool. For MCP servers, set it once in the server'smcp_toolsetconfig. - Claude Code: on by default.
- Cursor: dynamic context discovery, described above.
- VS Code: virtual tools, controlled by the
github.copilot.chat.virtualTools.thresholdsetting. GitHub describes grouping similar tools under one virtual tool that the agent expands when needed, and using embeddings to pre-select likely tools.
Anthropic suggests tool search once you have 10 or more tools, or more than 10,000 tokens of definitions.
3. Use a curated set for daily work
A curated set covers the common jobs in a few dozen tools. Keep it loaded, and reach into the full catalog only when a task needs an unusual endpoint. Anthropic gives the same advice for its API: keep your three to five most-used tools loaded and defer the rest.
4. Keep tool results small
- Ask a narrow question. "Orders over $500 this week" returns fewer rows than "show me my orders".
- Use limits and pagination when the tool offers them.
- Prefer a report to a raw export. In QuickBooks, ask for a profit and loss report through
quickbooks_run_reportbefore you pull every invoice.
5. If you build a server, keep the tool list stable
The 2026-07-28 MCP spec says servers should return tools/list in a deterministic order. A stable order lets clients cache the list, and it raises prompt cache hit rates when the tools sit in the model's context.
What PopMCP token trimming does
On PopMCP, token trimming means two things. A curated tool set is loaded for each provider you connect. The rest of that provider's catalog stays out of context and reachable through helper tools.
| Curated set | Full catalog | |
|---|---|---|
| What it is | The everyday tools for a provider. Shopify has 28, QuickBooks 28, HubSpot 55. | Every operation PopMCP has for that provider. Shopify has 70, QuickBooks 133, HubSpot 1,000+. |
| How the model reaches it | Loaded directly | Search, describe, then run |
| Plans | Every plan | Every plan. Free is read-only, so write tools are not exposed there. |
Every connected provider gets the same four helpers. For Shopify they are:
shopify_search_toolssearches the full catalog.shopify_describe_toolreturns the full schema for one tool.shopify_raw_operationruns an operation that is outside the curated set.shopify_connection_profilereports the connection's identity, status and access mode.
One account-level tool, list_connections, shows what the client was granted.
Take a refund. shopify_create_order_refund is not in Shopify's curated set. The model finds it with shopify_search_tools, reads its inputs with shopify_describe_tool, and runs it through shopify_raw_operation. On a paid plan a refund runs as soon as the model calls it, so try a read first. See MCP security best practices for the confirmation prompts your AI client offers.
The limits of this approach:
- It adds steps. A preloaded tool takes one call. A catalog tool takes two or three.
- PopMCP publishes no token-saving figure for it, and this post does not estimate one.
- If you use Claude Code, or the Claude API with tool search, the client already defers MCP tools from any server. You do not need PopMCP for that part.
- On Scale and Enterprise, the owner can load the full catalog up front for a chosen teammate and connection. That spends context, so use it sparingly.
The helpers are ordinary MCP tools. They work in any client that supports remote Streamable HTTP MCP with OAuth sign-in, including clients with no tool search of their own. PopMCP never meters tool calls, so the extra search call adds nothing to your PopMCP bill. The MCP server catalog shows the tool list for each provider.
Frequently asked questions
How many MCP tools is too many?
There is no single number. Anthropic's docs say tool selection degrades past 30 to 50 available tools, and they suggest tool search from 10 tools upward. VS Code stops a request outright at 128 enabled tools.
How do I fix "Cannot have more than 128 tools per request" in VS Code?
Open the tools picker in the Chat view and deselect tools or whole MCP servers until you are under 128. The VS Code docs also suggest enabling virtual tools through the github.copilot.chat.virtualTools.threshold setting.
Does Cursor have an MCP tool limit?
Cursor's MCP docs state no tool cap as of 5 October 2026. Since January 2026, Cursor gives the agent tool names up front and has it look up full descriptions when a task calls for them.
Why does my AI say the context window is full?
Instructions, tool definitions, chat history and tool results all share one window. Long chats, many loaded tools and large tool outputs all eat into it. Start a new chat, disconnect servers you are not using, and ask for filtered data.
What is MCP tool search?
Tool search replaces a long tool list with one search tool. The model searches the catalog, and only the matching definitions are loaded into context. Claude Code turns it on by default, and the Claude API offers it through defer_loading.
Do MCP tool definitions cost money?
Yes, when you pay per token. Anthropic's docs confirm that each definition loaded into context counts as input tokens, like any other tool definition. Deferring the tools a request does not use is the direct way to cut that.
Is a curated tool set enough?
For daily work, usually. Shopify's curated set on PopMCP covers orders, products, customers and inventory in 28 tools. You need the full catalog for jobs outside it, such as issuing a refund or creating a collection.
Sources
10 references, checked 5 October 2026
- Anthropic, Introducing advanced tool use (24 November 2025)anthropic.com
- Claude docs, Tool search toolplatform.claude.com
- Claude Code docs, MCP (tool search default, output limits)code.claude.com
- VS Code docs, agent tools (128-tool limit, virtual tools)code.visualstudio.com
- GitHub Blog, How we're making GitHub Copilot smarter with fewer tools (19 November 2025)github.blog
- Cursor, Dynamic context discovery (6 January 2026)cursor.com
- Cursor MCP docscursor.com
- Claude help center, custom connectors using remote MCPsupport.claude.com
- MCP 2026-07-28 changelog (deterministic tool order)modelcontextprotocol.io
- MCP tools specification (revision 2026-07-28)modelcontextprotocol.io