AI for investigators
Which AI Should an Investigator Use?
They aren’t interchangeable, and the differences that matter for casework aren’t the ones the benchmarks measure. Here’s what each one is actually good at, what it gets wrong, and where I use which.
- Reasoning and documents
- Claude
- Live X / real-time
- Grok
- Citations you can click
- Perplexity
- Never paste
- A subject’s PII
Read this part first. Which model you pick matters far less than what you put into it. Don’t paste a subject’s name, date of birth, Social Security number, address, phone number, account details or your client’s file into a consumer AI chat. Consumer tiers can retain conversations and may use them for training. You have no control over where that goes and no way to get it back.
If you handle data under GLBA, DPPA or FCRA permissible purpose, putting it in a chat box is a disclosure you didn’t have permission to make. Use it on your work — drafting, summarizing public documents, explaining a statute, writing a script — and keep the subject out of the prompt.
Claude
Best at: reasoning through something messy, long documents, writing that doesn’t sound like a press release, and anything with code in it.
If you hand it a 200-page bankruptcy schedule or a deposition transcript and ask what matters, it holds the whole thing in view and answers about the document rather than about documents in general. It’s also the one most likely to say it isn’t sure, which sounds like a weakness and is the opposite — a model that hedges when it should hedge is a model you can calibrate against.
Weak at: anything needing live data. It searches the web but it isn’t built around search the way Perplexity is, and it has no access to X at all. It also refuses more than the others, which is occasionally justified and occasionally just friction when you’re doing legitimate work on an unpleasant subject.
I use it for: programming, logic, and anything where being wrong would matter.
Grok
Best at: real-time, and X specifically.
This is the one genuinely unique capability in the list. Grok is wired into X, so it can search and summarize posts as they happen. For public records and business research it’s surprisingly strong, and the X connection means you can ask about a company or a person and get what’s being said right now rather than what was true at a training cutoff.
For OSINT that matters. An account’s posting history, who is talking about a business, what happened at an address last night — that’s a live source no other model reaches natively.
Weak at: the same thing that makes it useful. X is a sewer of unverified claims, and a summary of what X says is a summary of what X says — not a finding. It’s also less careful than Claude about distinguishing what it read from what it inferred, and it will state things with more confidence than the underlying posts support.
Treat every Grok answer as a lead. It tells you where to look. It doesn’t tell you what’s true.
ChatGPT
Best at: being a good all-rounder, and conversation. If you want to think out loud about a case, work through an approach, or have something explained until it lands, this is the most natural one to talk to. The ecosystem around it’s also the largest — more integrations, more tooling, more people who can help when you get stuck.
Weak at: telling you when it doesn’t know. It has a habit of agreeing with you, and of producing a confident, well-structured, completely wrong answer. The structure is the danger: a numbered list with headings reads as authoritative whether or not anything in it’s real.
I use it for: conversation and thinking through problems.
Perplexity
Best at: showing its work. Perplexity is built around search and it puts citations next to claims, so you can click through and read the source yourself.
For an investigator that’s the right shape. You should never be repeating what an AI told you; you should be reading what it found. Perplexity makes that the default rather than something you have to ask for.
Weak at: depth. It’s a research front-end more than a reasoning engine, and on a hard analytical question it will give you a competent summary where Claude would give you an argument. It also cites the web, which means it cites whatever ranks — a well-optimized content farm gets cited as readily as a court opinion.
Gemini
Best at: Google. If your work already lives in Gmail, Drive and Docs, having the model inside them removes a lot of copying and pasting. It handles very long inputs well and it’s strong on video and image.
Weak at: consistency. Quality varies more between questions than the others, and it’s the most likely to give you a different answer to the same question asked twice.
The open models
Llama, Mistral, DeepSeek, Qwen and the rest can be run on your own hardware. For an investigator there’s exactly one reason to care about that, and it’s a good one: nothing leaves your machine.
A model running locally on case material isn’t a disclosure. If you genuinely need to summarize a file containing subject data, this is the only responsible way to do it.
Weak at: everything else, relatively. A model small enough to run on a laptop is meaningfully worse than the frontier ones. And note that using DeepSeek’s hosted service isn’t local — that’s somebody else’s server, subject to somebody else’s jurisdiction.
What none of them can do
No model verifies anything. It produces text that resembles a correct answer, and on most questions that’s a correct answer. On the rest it’s a fluent, plausible, invented one, and it won’t sound any different.
The specific failure that has ended careers is the invented citation. Lawyers have been sanctioned for filing briefs quoting cases that never existed. The same failure mode applies to a case number, a statute section, a docket entry, an agency name or a fee. It will give you a case number in the right format for that county and the case won’t exist.
Every fact goes back to the source before it goes in a report. That rule doesn’t change based on which model you used or how confident it sounded. See what AI gets wrong for the specific failure modes and how to catch them.
A working setup
You don’t need all of them. Two covers most people:
- One strong reasoning model for documents, analysis and writing
- One live-data model for anything happening now
Add Perplexity if you want citations as a default rather than a request, and a local model if you regularly need to process material you can’t send anywhere.
The tiers move constantly — prices, limits and model names change every few months, and whatever is best this quarter may not be next. What doesn’t change is the division of labor: reasoning, live data, citations, privacy. Pick for the job, not for the brand.
Course
AI Essentials for Investigators
Where AI genuinely saves hours on a case, where it invents things that will end up in your report, and how to tell the difference before it matters.
See the course