Developers

Structured Output and Function Calling: Common Failure Modes

Function calling is reliable enough that it's easy to stop checking it carefully. The failures that do happen tend to be subtle, wrong-but-valid outputs rather than obvious crashes.

A&

AI & Tech Insights Team

September 30, 2026 · 3 min read

Modern models are reliable enough at structured output and function calling that it's tempting to stop scrutinizing it closely once a first integration works. The failures that do still happen tend to be the subtle kind, technically valid output that's still wrong, not the loud kind that throws an obvious error.

Schema-valid but semantically wrong

A model can return arguments that pass schema validation, correct types, all required fields present, while still being the wrong values for the actual situation, an amount field with a plausible-looking number that doesn't match what the user actually asked for, a date field with a valid date that's not the one implied by context. Schema validation alone catches structural problems, not semantic ones, and treating a passed validation as proof of a correct call misses this entire category.

Silently dropped optional fields

Optional parameters get omitted more often than expected, especially in longer, more complex tool schemas, without the model necessarily flagging any uncertainty about the omission. If an optional field actually mattered for correct behavior in a specific case, its silent absence can produce a technically valid but functionally incomplete call.

Hallucinated enum values

For fields constrained to a specific set of allowed values, a model can occasionally return a value close to but not exactly matching an allowed option, a plausible-sounding variant it effectively invented. Strict schema validation catches this as an error, which is the safer failure mode, but a looser validation approach can let a near-miss value through into downstream logic that assumed only the defined enum values were possible.

Wrong tool selected among similar options

When multiple available tools have overlapping purposes or similarly worded descriptions, a model can select a plausible but incorrect tool for a given request, one that superficially fits but isn't actually the right one for the specific case. This tends to get worse as the number of available tools grows and their descriptions become less sharply distinguished from each other.

Retrying with the same mistake

When a tool call fails and gets fed back to the model for retry, a model without enough context about why it failed can retry with a similar or identical mistake, rather than genuinely correcting course, especially if the error message returned wasn't specific enough to make the actual problem clear.

What actually reduces these failures

Strict schema validation as a hard gate, rejecting any output that doesn't conform rather than trying to coerce a near-miss into something usable. Clear, sharply distinguished tool descriptions and names, reducing ambiguity between similar tools rather than relying on the model to infer subtle intended differences. Specific, actionable error messages fed back on failure, so a retry has real information to correct against rather than just "that didn't work, try again." And logging every tool call with its arguments for review, since a lot of these failures are invisible unless someone's actually looking at the call history, not just the final task outcome.

The realistic takeaway

Function calling reliability has genuinely improved, but "reliable most of the time" still means a real, non-trivial failure rate at any meaningful production volume, and the failures that do occur tend to be the quiet, semantically wrong kind that a pure schema check won't catch on its own.

Share:XLinkedInWhatsApp

© 2026 AI & Tech Insights. All rights reserved. This article may not be reproduced without permission. See our disclaimer.

Related articles

Get new guides by email

Useful AI and tech guides, occasionally. No unnecessary emails.