A radiologist in Houston named Keith Dreyer ran a study in 2023 comparing AI-assisted reads on chest X-rays to unassisted ones. The AI — a system built on Google's Med-PaLM 2 — didn't outperform the radiologists. What it did was catch the cases the radiologists were most likely to miss when fatigued. That asymmetry is the whole story of where AI tools are actually useful: not replacing competence, but plugging the gaps competence leaves.
This article is a tour of the tools that are genuinely doing that — in writing, coding, research, image and video creation, and knowledge work broadly. For each one, you'll get what it's concretely good at, where it falls apart, and who shouldn't bother. No inflated promises. The AI industry generates enough of those without our help.
ChatGPT and Claude: Where the Reasoning Actually Lives
OpenAI's ChatGPT, specifically the GPT-4o model, and Anthropic's Claude 3.5 Sonnet are the two large language models most worth knowing in depth. They're similar in many surface ways — both are conversational, both handle long documents, both write code — but they differ in ways that matter depending on what you're doing.
ChatGPT is faster at structured output. If you need a table, a JSON object, a formatted report, or a step-by-step plan with numbered sub-steps, GPT-4o produces it more reliably than almost anything else. Its integration with tools like the Code Interpreter (now called Advanced Data Analysis) means you can drop in a spreadsheet and ask it to spot anomalies — a genuinely useful trick for anyone dealing with messy data who doesn't want to write Python.
Claude 3.5 Sonnet, on the other hand, is the better writer. Its prose sounds less like it was generated by a model that trained on Reddit and Wikipedia summaries. For legal drafts, nuanced emails, technical documentation that has to be clear without being condescending, Claude's output requires less editing. It also handles very long context windows — up to 200,000 tokens as of late 2024 — meaning you can feed it an entire book manuscript and ask coherent questions about it.
Where both fail: they hallucinate facts with confidence. Neither should be trusted to report specific statistics, legal citations, or historical dates without verification. If you treat them as draft machines rather than oracles, they're extraordinary. If you treat them as encyclopedias, you will be embarrassed at some point.
GitHub Copilot and Cursor: The Coding Tools With Real Productivity Data
In 2022, GitHub published a controlled study (conducted with researchers at MIT) showing that developers using Copilot completed tasks 55% faster than those who didn't. That number has been disputed and picked apart — the tasks were narrow, the timeframe was short — but the general direction has held up in subsequent workplace studies. Copilot makes programmers faster. The question is: at what?
Copilot is excellent at boilerplate. If you're writing a new REST API endpoint in Python, or scaffolding a React component, or doing anything you've done a dozen times in a slightly different form, Copilot finishes sentences in a way that feels almost telepathic. It's trained on enough open-source code that it can produce functional implementations of common patterns with almost no prompting.
Where it becomes genuinely impressive is in unfamiliar territory — specifically, writing code in a language or framework you're learning. A JavaScript developer picking up Rust will find Copilot dramatically more useful than a Rust developer, because the Rust developer already knows what to type; the JS developer would have spent 20 minutes reading Stack Overflow instead.
Cursor is a newer entrant, built as a fork of VS Code that puts AI at the center rather than treating it as a plugin. Its standout feature is the ability to select code, describe what's wrong with it in plain language, and have the AI rewrite it in place. The context it maintains across files in a project is better than Copilot's, which matters for larger codebases. Several engineering teams at mid-size companies have publicly switched their default editor to Cursor over the past year. It's worth evaluating seriously if you write code all day.
Who should skip both: if you write code occasionally for simple scripts and already know what you're doing, these tools add friction more than they remove it. They're most transformative for people who code professionally, or for non-developers trying to produce working code without a deep background.
Perplexity AI: Research That Cites Its Sources
The core problem with using ChatGPT for research is that it doesn't tell you where it got anything. You get an answer, it sounds authoritative, and then you have to go fact-check it from scratch — which defeats a large part of the purpose. Perplexity AI solves this by building a search engine on top of a language model. Every claim it makes is linked to a source you can click.
In practice, Perplexity works best for questions that have good web coverage. Ask it about the current state of mRNA vaccine research, and you'll get a clean summary with links to Nature, The Lancet, and relevant preprints. Ask it something narrow and specialized — say, the internal budget disagreements within a specific municipal government — and it will either admit it doesn't know or construct something plausible from thin sources. The source citations are only as good as the underlying web coverage.
Its 'Pro Search' mode, which requires a subscription, can run multi-step searches — essentially breaking a complex question into sub-questions, retrieving sources for each, then synthesizing. This is useful for competitive research, background reading before a meeting, or getting up to speed on an unfamiliar industry. A journalist friend of mine uses it as a first-pass research tool and estimates it cuts her background research time by about a third. That's a reasonable benchmark.
The honest limitation: Perplexity still gets things wrong, especially about recent events where web coverage is contradictory. It also tends to surface whichever sources happen to rank well in search, not necessarily the most authoritative ones. Use it to find the questions you should be asking and the sources you should be reading — not as the last word on anything.
Midjourney and Stable Diffusion: Image Generation for Different Types of Users
Midjourney v6, released in early 2024, produces images that regularly fool people who aren't looking carefully. The photorealistic outputs — particularly portraits, landscapes, and architectural renders — have a quality that would have been implausible from a text-to-image system two years ago. For designers doing mood boards, marketers building campaigns without a photography budget, or anyone who needs a compelling visual fast, it's become a serious professional tool.
The interface is the friction point. Midjourney still runs primarily through Discord, which means joining a server, using slash commands, and wading through other people's generations in public channels (unless you pay for private mode). This is genuinely annoying and represents a real usability gap compared to what you'd expect from a tool at its price point. Rumors of a standalone web app have circulated for over a year; by the time you read this, it may finally exist.
Stable Diffusion is the open-source alternative. It runs locally on your machine (if your GPU is powerful enough — you'll want at least an NVIDIA RTX 3080 or equivalent), costs nothing after setup, and can be fine-tuned on specific image styles or faces using techniques like DreamBooth and LoRA. The ceiling is higher than Midjourney's for specialized use. The floor is lower — default outputs without fine-tuning or careful prompting are noticeably rougher. It's the right tool for developers, researchers, and anyone who needs control over the model itself, not just its outputs.
Adobe Firefly deserves a mention here because it's the only major image AI trained on licensed content, which matters enormously for commercial use. If you're producing images for advertising or publishing and need indemnification against copyright claims, Firefly is the safer choice, even if the raw quality ceiling is slightly below Midjourney's.
Runway and Sora: The Video Tools That Are Still Finding Their Feet
OpenAI's Sora generated a wave of justified amazement when its demo clips surfaced in February 2024. The outputs — people walking through Tokyo, historical footage that never existed, abstract simulations — looked like nothing a text-to-video system had produced before. By the time of its actual public release in late 2024, the reality was slightly more modest: it's an extraordinary tool for short clips and visual ideation, less reliable for anything requiring consistent characters or long sequences.
Runway Gen-3 Alpha is the working professional's choice right now. It's available, it's fast, and its 'Act-One' feature — which animates a character based on a real actor's performance captured on a simple camera — is genuinely novel. Film students and indie directors have been using it to storyboard sequences and create rough animatics that would have required expensive motion capture setups a few years ago.
The honest state of video AI as of late 2024: it's impressive for five-to-ten second clips, useful for visual concept work, and not yet reliable for narrative filmmaking. Character consistency across scenes — the same person looking the same way in different shots — remains a significant unsolved problem. Temporal coherence (things moving the way physics says they should) is better than it was but still occasionally produces the uncanny distortions that mark AI video to a trained eye.
If you're in advertising or social media, these tools are worth learning now. If you're in long-form filmmaking, watch the space closely but don't restructure your production pipeline around it yet.
The Boring but Critical AI Tools: Otter, Notion AI, and the Productivity Layer
The highest-ROI AI tools for most office workers aren't the glamorous ones. They're meeting transcription, document summarization, and writing assistance baked into tools you already use.
Otter.ai transcribes meetings in real time, labels speakers, and produces a summary with action items. The transcription accuracy is around 90-95% for clear audio with native English speakers — good enough to replace manual note-taking entirely. Where it degrades: heavy accents, technical jargon, crosstalk, and anyone who speaks quietly. In those situations you'll spend more time correcting the transcript than you saved by generating it. But for a standard team standup or client call, it works well enough that not using something like it feels like a deliberate choice to waste time.
Notion AI, embedded inside Notion's workspace, is a quieter story. It can summarize a long document, extract action items from a meeting note, rewrite a paragraph in a different tone, or translate a page — all without leaving the document you're working in. The quality is slightly behind standalone Claude or GPT-4o, but the friction is so much lower that for quick tasks it wins on pure convenience.
Microsoft Copilot for Microsoft 365 — the version that integrates with Word, Excel, PowerPoint, and Teams — deserves attention simply because of scale. If your organization runs on Microsoft tools, it can draft emails from bullet points in Outlook, generate slides from a document in PowerPoint, and surface relevant information from across your SharePoint in a chat interface. The quality is uneven across applications (it's best in Word, shakiest in PowerPoint), and the pricing is steep for small organizations. But for enterprises already deep in the Microsoft ecosystem, it's the most natural AI integration available.
The general point here is important: AI tools that live inside your existing workflow will beat better standalone tools for routine tasks, almost every time. The friction of switching contexts is a real cost that doesn't show up in feature comparisons.
What to Actually Watch: Where This Goes in the Next Few Years
Multimodal reasoning is where the frontier is moving fastest. GPT-4o can already look at an image and reason about it alongside text. Google's Gemini 1.5 Pro can process an hour of video as context. The near-term trajectory is models that understand whatever format information arrives in — text, image, audio, video, code — and reason across all of them simultaneously. The applications that follow from that are hard to fully predict, but medical imaging, scientific research, and engineering will be transformed before most consumer applications are.
Agents — AI systems that don't just answer questions but take sequences of actions autonomously — are the other major frontier. OpenAI's Operator, Anthropic's computer-use features in Claude, and Google's Project Mariner are all early attempts to let AI models browse the web, fill out forms, and complete multi-step tasks without human intervention at each step. Current reliability is low enough that you'd be unwise to deploy agents for anything consequential without a human in the loop. But the improvement curve on this is steep, and within two to three years, autonomous agents handling routine knowledge work tasks is not a fanciful prediction — it's the disclosed roadmap of every major lab.
The useful framing for anyone trying to stay oriented: ask not which AI tool is most impressive, but which reduces the cost of something you already know is valuable. The radiologist example at the start of this article is a good model. The AI wasn't better at radiology. It was good at the specific failure mode — fatigue-related misses — that human radiologists couldn't easily fix themselves. Find the equivalent in your own work, and you'll find where AI is actually worth your time.
Frequently Asked Questions
What is the best AI tool for writing right now?
Claude 3.5 Sonnet (from Anthropic) produces the best prose of any current AI model — the output requires less editing and sounds less obviously machine-generated. ChatGPT with GPT-4o is a strong second, particularly for structured documents. If you're writing for publication or producing anything where voice and clarity matter, start with Claude.
Is GitHub Copilot actually worth it for professional developers?
Yes, for most professional developers. A 2022 MIT/GitHub study found a 55% speed improvement on coding tasks, and while that figure applies to specific conditions, the general productivity gain has held up in workplace evaluations. It's most valuable for boilerplate, unfamiliar frameworks, and developers learning new languages. It's less transformative for experienced engineers working in deeply familiar codebases.
Can AI tools be trusted for factual research?
Not without verification. ChatGPT and Claude will state incorrect facts with the same confidence as correct ones — this is a structural property of how they work, not a fixable bug. Perplexity AI is more reliable for research because it cites sources you can check, but those sources are only as good as web coverage on the topic. Treat any AI-generated factual claim as a starting point for verification, not an endpoint.
What AI image generator should I use for commercial work?
Adobe Firefly is the safest choice for commercial use because it was trained on licensed content and Adobe offers indemnification for commercial outputs. Midjourney produces higher-quality images for general creative work but its licensing terms for commercial use are less clear-cut. If legal certainty matters — advertising, publishing, branded content — Firefly is the right call even if the raw output quality ceiling is slightly lower.
What's the difference between ChatGPT and Claude?
GPT-4o (ChatGPT) is better at structured outputs, tool use, data analysis, and working with plugins and integrations. Claude 3.5 Sonnet is better at writing prose that doesn't need heavy editing, handling very long documents (up to 200,000 tokens), and nuanced tasks where the quality of the language itself matters. Most power users keep both and switch based on the task.
Are AI video tools like Sora and Runway ready for professional filmmaking?
Not for long-form narrative filmmaking — character consistency across scenes and realistic physics in long sequences remain unsolved problems. For short clips, visual concept work, storyboarding, and social media content, Runway Gen-3 Alpha is a practical tool right now. Revisit Sora and the broader category in 12-18 months; the improvement rate in this area is fast.
What AI tools are most useful for people who don't write code?
Perplexity for research, Claude or ChatGPT for writing and document work, Otter.ai for meeting transcription, and Midjourney for images. If you're inside a Microsoft organization, Copilot for Microsoft 365 is worth trying for email drafting and document summarization. These tools add real value without requiring any technical background.
Will AI agents replace knowledge workers in the near future?
Current AI agents — including OpenAI's Operator and Anthropic's computer-use features — are unreliable enough that deploying them autonomously on consequential tasks would be a mistake. They make errors, get stuck in loops, and require supervision. The honest answer is that within two to five years, agents will handle many routine knowledge work tasks reliably, but that prediction depends on improvement rates continuing at their current pace, which is not guaranteed.