{"slug":"reliable-json-models-for-agents","question":"Which LLM reliably returns valid structured output for agents?","title_zh":"Agent 需要稳定 JSON 时该选哪个模型？","answer":"There is no universally reliable JSON model: reliability changes with the schema and the application parser. Test the exact schema more than once, then compare pass rate before cost. AI NetCafé publishes live-app benchmarks where a single format failure makes the run fail.","answer_zh":"不存在对所有 JSON 都最稳的模型。可靠性会随 schema 和应用解析器变化，必须用真实格式重复测试，先看通过率，再比较成本。","updated":"2026-09-16T20:30:05.017Z","method":"Live application tasks are run repeatedly; pass/fail, actual billed USD cost, and latency are recorded. Upstream-unavailable models are excluded from quality conclusions.","tool":{"name":"compare_models","description":"Run one prompt across multiple LLMs in parallel and return every answer side by side with measured platform cost metadata and latency. The beta platform covers the user charge ($0.00). This answers \"which model should I actually use for this kind of task?\" with data instead of guesswork. Example — GET https://ainetcafe.com/t/compare_models?prompt=Explain+CAP+theorem+in+1+line","arguments_example":{"prompt":"Return only valid JSON matching this schema: {\"title\":\"string\",\"risks\":[\"string\"]}."},"mcp_endpoint":"https://ainetcafe.com/mcp","openapi":"https://ainetcafe.com/openapi.json"},"steps":["Paste the exact schema and forbid prose outside JSON.","Run each candidate more than once.","Treat any parse failure as a failed production run.","Choose the cheapest model among those with a complete pass rate."],"evidence":[{"slug":"gpt-researcher","available":true,"url":"https://ainetcafe.com/bench/gpt-researcher","data_url":"https://ainetcafe.com/api/bench/gpt-researcher","updated":"2026-09-16T20:18:37.711Z","runs_per_task":2,"models_tested":3,"models_with_complete_pass_rate":3,"best_measured_value":{"model":"gpt-5.6-luna","usd_per_run":0},"failed_all_tasks":[],"not_measured":[]},{"slug":"presenton","available":true,"url":"https://ainetcafe.com/bench/presenton","data_url":"https://ainetcafe.com/api/bench/presenton","updated":"2026-09-16T20:30:05.017Z","runs_per_task":2,"models_tested":3,"models_with_complete_pass_rate":3,"best_measured_value":{"model":"gpt-5.6-luna","usd_per_run":0},"failed_all_tasks":[],"not_measured":[]}],"page":"https://ainetcafe.com/agent-guides/reliable-json-models-for-agents","machine_readable":"https://ainetcafe.com/api/agent-guides/reliable-json-models-for-agents","license":"CC BY 4.0"}