The vault_drop tool disappeared from ChatGPT the moment I labeled it honestly.
Nothing had failed at the transport layer. The client could reach the same MCP endpoint that claude.ai and Manus were using. The tool server was healthy. The tool description was present. But ChatGPT’s runtime saw a write annotation and hard-gated the tool out of chat and voice.
That left me with an awkward fact: a tool that writes a file is not automatically dangerous, and a safety label can be wrong about the operation that matters. I had assumed the generic, technically accurate label was the responsible choice, the same annotation everywhere, no client-specific carve-outs. I did not think about what a runtime I don’t control would do with that label until a real feature quietly stopped working for anyone using it through ChatGPT, and I had to trace an absence back to metadata before I understood why.
The thing the annotation missed
vault_drop does one job. It creates a new Markdown file in a sandboxed inbox directory. It cannot overwrite an existing file. It cannot edit a file in place. It cannot choose an arbitrary destination. Its payload is capped at 256 KB.
That is a write in the literal sense. A new file exists after the call. So the honest MCP annotation is write-capable, and that is what claude.ai continues to receive.
Writing bytes is not the relevant risk. vault_drop cannot alter existing records, escape its directory, or hide an unrestricted operation behind a friendly name.
ChatGPT did not make that distinction. Its runtime treated the write annotation as a hard availability gate. The tool did not become a confirmation step or a visibly restricted action. It became unavailable. That was the diagnostic path: first I checked the endpoint and tool registration, then I traced the absence to the annotation gate. What looked like a connector problem was client policy acting exactly as designed.
One tool, two advertised contracts
The change landed in server/mcp.ts as commit commit-4C2A: 10 additions and 1 deletion.
server/mcp.ts | 11 ++++++++++-1 file changed, 10 insertions(+), 1 deletion(-)I did not fork the tool. I did not create a second MCP endpoint. I did not loosen the sandbox, raise the 256 KB cap, or grant overwrite access. I changed the metadata advertised for the ChatGPT-facing client. It receives vault_drop as read-only. claude.ai keeps the write annotation, because that is still the accurate generic classification for a tool that creates a file.
That is the lie in the title. The endpoint is not lying about the operation. It is lying about the risk category needed to get past one client’s coarse safety gate.
The implementation is client-specific only at advertisement time. All three clients still invoke the same tool with the same two inputs:
vault_drop(content, filename)The server applies the same new-file-only behavior after the call, regardless of which client supplied it. The divergent annotation changes discoverability. It does not change authority.
Why I did not solve this with another tool
Leaving the write annotation everywhere was the cleanest semantic choice, and it lost because it removed the use case in ChatGPT chat and voice. A perfect label is not useful if the client turns it into a permanent off switch.
A second tool with a more comforting name would have been worse. It would duplicate the contract, create two descriptions to maintain, and invite drift between the supposedly safe tool and the original. The next feature request would create the same argument all over again: which of the two tools gets it, and do their constraints still match?
I also did not want the client to infer safety from prose alone. A description that says “new files only” does not help if the runtime evaluates the annotation before a person or model can act on that detail. The enforcement has to stay in the server behavior, where the sandboxed directory, no-overwrite rule, and size limit are real constraints rather than promises in a tool description.
The annotation is part of the product surface
The same tool implementation can be reachable, hidden, blocked, or framed as safe depending on how a client interprets one field.
That means the annotation deserves the same review as the implementation. I now ask two separate questions when wiring a tool into a new client: what can this tool actually do, and what will this client allow the user to do once it sees that label?
Those answers matched for claude.ai. They did not match for ChatGPT. Keeping the generic label would have made a constrained tool unusable.
So I kept the constraint real, kept the endpoint shared, and changed the label where the runtime could not see the difference. The tool is still new-file-only. The safety boundary is still on the server. The lie is limited to the one place where a literal label prevented the safer behavior from being used at all.