Chatbots Are Accurate but Rarely Point to Official Sources, Microsoft Copilot Adds a Prompt-to-App Builder, and Gartner Expects Many Agent Projects to Be Cancelled
A States United study found ChatGPT and Google AI cut factual errors on voter questions but linked to state election sites less than half the time, Microsoft's revamped Copilot adds a natural-language app builder called Code, and Gartner's forecast that over 40% of agentic AI projects will be cancelled by 2027 is a useful checklist for builders.
Three stories this week share a theme: accuracy is no longer the only thing to measure. Where an answer points, who is allowed to build on top of the assistant, and whether the project survives contact with cost and risk all decide if a chatbot succeeds in production.
1. Accurate Answers, Weak Pointers to Official Sources
A States United Democracy Center study of ChatGPT and Google AI, published ahead of the 2026 midterms, describes a split result. Per the study's summary, verifiable factual errors on voter questions fell from 6.9% (Google AI) and 8.2% (ChatGPT) in an early round to 0% in a later primary round. That is real progress.
The weaknesses sit around the facts. The study reports that the assistants sent voters to official state election websites less than half the time: ChatGPT mentioned a state election site in 39.4% of responses and Google AI in 55.6%. ChatGPT reportedly returned incomplete candidate lists for nearly 90% of gubernatorial race queries, and Wikipedia accounted for more than 12% of all cited links.
Two lessons carry beyond elections. First, an eval that only checks factual correctness will give you a green dashboard for an answer that is correct, incomplete and unsourced. Second, citation quality is a measurable product property. You can score it:
const OFFICIAL = new Set(['vote.gov', 'sos.state.gov', 'fec.gov']);
function citationScore(answer: Answer, domainAllowlist = OFFICIAL) {
const domains = answer.sources.map((s) => new URL(s.url).hostname);
return {
hasAnySource: domains.length > 0,
officialShare: domains.filter((d) => domainAllowlist.has(d)).length / Math.max(domains.length, 1),
completeness: coveredItems(answer.text, answer.expectedItems), // e.g. all candidates listed
};
}
Run it per topic and per language, and for high-stakes topics prefer a retrieval layer restricted to an allowlist of authoritative domains, with a retrieval date shown to the user. A table of official-source rate by topic would make a good dashboard visual.
2. Copilot Adds Code: Users Become Builders
On September 25 Microsoft unveiled a revamped Copilot app organized around Home, Code and Autopilot. Code lets users, including people without a development background, describe a small app, tracker, dashboard or automation in natural language and have it built. Coverage says early-access customers get it at the end of the month, with Microsoft 365 Premium and Pro subscribers previewing it later this year. We have not independently verified pricing or the underlying models.
For builders, the shift is that the chat box becomes a way to create artifacts that persist and run, not just text. That raises the stakes on ownership and review. Questions to settle before you offer something similar: who can see and edit what a user generates, what data connections a generated app inherits, how generated code is scanned, and how a non-developer rolls back a bad change. Treat a generated app like any other deployed code, with an owner, permissions and an audit trail.
3. Reading Gartner's Cancellation Forecast as a Checklist
Industry roundups this week keep citing the forecast that over 40% of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls. It is a prediction, not an outcome, and the same summaries cite projections that task-specific agents will appear in about 40% of enterprise applications by the end of 2026. Both can be true: many will start, and many will stall.
The three named causes map neatly onto things you can instrument from day one:
- Cost: log tokens and tool calls per completed task, not per message, and set a budget alert.
- Value: pick one outcome metric (resolved tickets, completed orders) before launch and compare against a human baseline.
- Risk: give every agent its own identity, least-privilege permissions and an approval step for irreversible actions.
Conclusion
Measure what surrounds the answer: where it points, what users can build from it, and what it costs and risks per task. A model can be accurate and still fail a user who needed the official source, the complete list or a project that survives its first budget review.
— Maya
Frequently asked questions
How accurate are chatbots on voting questions now?
A States United Democracy Center study reports verifiable factual errors fell to 0% for ChatGPT and Google AI in a later primary round, down from 6.9% and 8.2% earlier. The same study found the assistants linked to official state election sites less than half the time and often gave incomplete candidate lists, so confirm details with official election offices.
What is Copilot Code?
Code is a part of Microsoft's revamped Copilot app, announced September 25, 2026, that lets users build small apps, dashboards and automations by describing them in natural language. Reporting says early-access customers get it at the end of September, with Microsoft 365 Premium and Pro previews later this year.
Why might agentic AI projects be cancelled?
Gartner has predicted that over 40% of agentic AI projects will be cancelled by the end of 2027 due to escalating costs, unclear business value or inadequate risk controls. It is a forecast. Tracking cost per task, an outcome metric and scoped agent permissions addresses all three causes.
I'm Maya — I write most of what you'll read here. I spent years as a copywriter before I got a little obsessed with what these AI tools can actually do, so now I spend my days poking at chatbots, breaking them, and writing up what's worth your time. Everything here is something I've actually tried. If a prompt didn't work for me, it doesn't make the cut.
Want to try any of this?
Smillee's free and there's no signup — open it and paste in whatever you're working on.
Try it free: AI Image Generator →More from the blog
- Trends
Claude Sonnet 5.5 Held Its Price While Getting Faster, OpenAI Dots Skipped Europe and the UK at Launch, and Voters Are Now Using Chatbots as Ballot Guides
Claude Sonnet 5.5 shipped September 28 at unchanged $2/$10 pricing with a reported 30% speed gain, OpenAI's Dots agents launched without support for Pro users in the EEA, Switzerland and the UK, and midterm voters are leaning on chatbots whose election answers have been uneven. Three lessons about model upgrades, regional rollouts and high-stakes answers.
- Trends
GPT-6.1 Sol Matches Its Flagship at One-Fifth the Price, Gemini 4 Argon Ships to Cyber Defenders First, and Inworld Buys Ultravox to Own the Voice Stack
OpenAI released GPT-6.1 Sol at $2/$10 per million tokens with near-Astra coding scores, Google gated Gemini 4 Argon behind its Fairwind Program for trusted cyber defenders, and Inworld acquired voice-agent platform Ultravox. Three signals about pricing, release strategy and stack consolidation for chatbot builders.
- Trends
Microsoft Gave Its Copilot Agent a Directory Identity, Claude Went to FedRAMP High, and Approval Gates Became a Product Feature
Microsoft revamped Copilot with Code and an always-on Autopilot agent that carries its own directory identity, Anthropic brought Claude to government under FedRAMP High, and low-code platforms like UiPath added tool-call confirmations. Three signs that identity, compliance and human approval are becoming core chatbot infrastructure.