AI Chat Assistant Comparison Cheat Sheet
The feature axes that actually differ between assistants
Assistants converge on quality and diverge on everything around it. These are the axes worth checking before committing a team to one.
Axes to compare
| Axis | What to check |
|---|---|
| Context length | The real usable limit on your plan, not the headline number |
| File handling | Which formats, how many at once, and whether they persist between chats |
| Web access | Whether it browses live, and whether it cites what it read |
| Code execution | Whether code runs in a sandbox and what that sandbox can reach |
| Memory | Whether it remembers across sessions, and how to inspect and clear that |
| Integrations | Connectors to the tools your team actually uses |
| Data retention | Whether your inputs train the model, and the opt-out path |
| Admin controls | SSO, audit logs, per-seat policy — the reason procurement says no |
Evaluating honestly
- Test with YOUR documents, not a demo prompt Every assistant looks excellent on a clean example
- Write the twenty questions first, then run them on each Otherwise you grade on vibes and pick the one you tried last
- Include questions whose honest answer is "I do not know" How a tool handles absence matters more than how it handles ease
- Re-test after a month These products change under you between evaluations
Frequently asked questions
How should I actually compare assistants?
Write your twenty real questions first, then run all of them through each candidate. Grading afterwards on impressions reliably picks whichever one you tried most recently.
Does my data train the model?
It depends on the product and the plan, and it changes. Check the current terms for the exact tier you are buying and confirm the opt-out path before rolling it out to a team.
Was this cheat sheet useful?
Comments
No comments yet — be the first.
Need a different cheat sheet?
Tell us what you would like to see and we will build it — free.
Request a cheat sheet