I've been seeing a lot of questions lately about how different models handle long-context, complex tasks, so I decided to run a practical test. I uploaded a single 50-page software licensing agreement (PDF) to Claude 3 Opus, GPT-4 Turbo (128K), and Kimi (the 200K context version). My goal was to see how effectively each could analyze the document without any prior chunking or summarization from me.
I asked all three the same set of five questions: identify the governing law clause, list all payment milestones, summarize the IP assignment terms, highlight any unusual liability caps, and find the term renewal conditions. The key was to see if they could accurately locate and interpret details scattered throughout the full document.
Here's a quick rundown of what stood out:
* **Claude 3 Opus** provided extremely thorough, well-reasoned answers, almost like a careful paralegal. It excelled at interpreting nuanced clauses and connecting related terms. However, it was the slowest and most expensive run of the three.
* **GPT-4 Turbo** was very fast and confident in its answers, with a polished, direct style. It correctly identified most clauses but missed one specific payment detail buried in a schedule. Its answers felt more "executive summary."
* **Kimi** was notably the fastest and handled the 200K context without issue. Its answers were accurate for the most part, but phrased more concisely—sometimes I wished for a bit more interpretive depth on the liability language. The cost-effectiveness for this volume of text is its strong suit.
For community reviews, the takeaway seems to be that the "best" tool depends heavily on the *type* of analysis needed. For sheer speed and cost on a massive document, Kimi is compelling. For deep, nuanced interpretation where time and cost are less critical, Claude shines. GPT-4 Turbo sits in a confident middle ground.
Has anyone else done a similar side-by-side comparison for long-document analysis? I'm particularly curious if others have tested retrieval accuracy on specific numerical details or dates across different file formats.
Keep it real, keep it kind.
That's really helpful, thanks for running this test. I'm curious about the missed payment detail from GPT-4. Was it just a straight-up omission, or did it maybe misinterpret a conditional phrase or date? That kind of error is the tricky part for project timelines.
Also, on the speed vs. cost point for Claude, how would you decide when that level of detail is actually worth it? Like for a final contract review, sure, but maybe for a quick internal summary during negotiations, the faster option wins?