Learnings From 100+ Experiments Comparing LLMs for AI Coding

This ad is not shown to multipass and full ticket holders
React Summit US
React Summit US 2026
November 17 - 20, 2026
New York, US & Online
Upcoming event
React Summit US 2026
React Summit US 2026
November 17 - 20, 2026. New York, US & Online
Bookmark
Rate this content
Sentry
Promoted
Code breaks, fix it faster

Crashes, slowdowns, regressions in prod. Seer by Sentry unifies traces, replays, errors, profiles to find root causes fast.

Get started

On my YouTube channel AI Coding Daily, I've published 100+ videos comparing different models for coding: Opus vs GPT, Kimi vs GLM, New vs Older versions, Effort Medium vs High, etc. Now I see clear patterns for evaluating models and deciding which one to choose for specific tasks and projects.

This talk has been presented at AI Coding Summit London, check out the latest edition of this Tech Conference.

Povilas Korop
Povilas Korop
28 min
06 Jul, 2026

Comments

Sign in or register to post your comment.
Video Summary and Transcription
Introduction to LLMs from a Developer's Perspective, YouTuber with AI Coding Daily channel, evaluates new models and versions on YouTube, gaining traction and positive feedback. Exploring the Best Models: My Benchmark on 18 LLMs, evaluating cost, points, and competition among models for day-to-day use. Measuring Models: Prompt Methodology, Tests on PHP, Laravel, React, TypeScript, and CSV, with emphasis on correct data usage and benchmark awareness. Model Evaluation: Chinese models' progression, benchmark awareness, and Opus and GPT superiority. Model Pricing and Quality Comparison: Opus, GPT, and Composer 2.5 cost-effective options. Composer 2.5 and GPT 5.4 mini offer quality at low prices. Choosing models based on project needs and budget constraints. Discussion on local LLM investment challenges, SONnet 5 benchmark performance, and exploration of new models like Proton Luma 2.0. Importance of skill definitions for model performance and limitations in local model implementation. Discussion on the importance of model harness, setup considerations, and evaluation methodologies for LLMs.

1. Introduction to LLMs and AI Coding Daily

Short description:

Introduction to LLMs from a Developer's Perspective, YouTuber with AI Coding Daily channel, evaluates new models and versions on YouTube, gaining traction and positive feedback.

Hello, guys. Glad to be here. Honored to have such an audience. So the amount of questions to Mario has actually freaked me out a bit, so this will be kind of the main part. But before the main part starts, I'll talk about LLMs. Mario's kind of depicted the Skynet scenario here, right? So everything is going to be worse. And now I will try to bring us back to the ground to the daily tasks of individual developers.

How do they use LLMs and the amount of experiments I did around that from an individual developer perspective, not that much from an organizational point of view. So this is briefly about me. I'm a YouTuber. My 10-year-old daughter is proud to tell that to her friends that dad is a YouTuber. But I'm a developer with 20 years of experience, mostly about PHP and Laravel framework. And I started creating content 10 years ago and basically never stopped.

So one of my channels, which I have four now, is AI Coding Daily. And this is exactly what I will talk about. And this is what I do on my YouTube channel. I evaluate models. I do a lot of other topics as well. But this kind of got traction. So whenever some new model or version comes up, I test it on my own. Benchmark is probably too strong a word because I'm not an organization. I cannot compete with SWE Bench or other official benchmarks. But I run the models with my prompt and my automatic evaluation from the point of evaluating the code.

So whenever some new model comes up on Twitter, you see those flappy birds and clones of various games and HTML experiments. And I personally hate those because they are trying to one-shot some project which will never go out as is anyway, which is not a realistic scenario. So I was trying to do basically this new model. I tried to evaluate, to run prompts, to run evaluations. And people started liking that. So it became one of my sexy topics basically on YouTube. And these are usual comments I get.

2. Exploring Model Competition and Benchmark

Short description:

Exploring the Best Models: My Benchmark on 18 LLMs, evaluating cost, points, and competition among models for day-to-day use.

So similar to the question, what is the safest model to use right now? People are always curious about what do we use in terms of tools and the models. And these are comments, just random screenshots, but there are many more, which makes it difficult now with so many models being good enough. And it's not just model. Can you test Opus 4.8 high? Because high is different than medium. And there's also X high. So you see the thing, there are so many models and so many toys for me to play around with. And that's how my leaderboard was born.

In total, I have over 100 videos on my YouTube channel. So this is the result that I will be talking about, my own benchmark, which models are basically the best. But it's not necessarily, I cannot, again, I cannot compete with SWBench. So it's for my use cases, which will show my methodology and how I evaluate in a minute. But you see, basically, it's good that the screen is pretty big and I will zoom that in bit by bit a bit later. But I have 18 LLMs, score max of 25 points on five projects. And I also evaluate the average time per prompt until the end and average cost per prompt, which is API pricing.

And this is the same table, but in cost and points. So points max up to 25 on the left and the cost is to the right, which is not really the perfect way to describe that visually. I guess it could be skewed or it could be like zoomed in or out. But what I want to emphasize with this is like this is like from cheap to expensive and from bad to better. And this is basically the best place to be in. And you see how big is the competition. Many more models became good enough in terms of cost and value for, again, day to day use, that the competition is pretty fierce. Many developers, many people swear by their model, they use and they're all basically right. A lot of models are good enough these days. So this is basically so that you understand the ecosystem.

QnA

Check out more articles and videos

We constantly think of articles and videos that might spark Git people interest / skill us up or help building a stellar career

Don't Solve Problems, Eliminate Them
React Advanced 2021React Advanced 2021
39 min
Don't Solve Problems, Eliminate Them
Top Content
Kent C. Dodds discusses the concept of problem elimination rather than just problem-solving. He introduces the idea of a problem tree and the importance of avoiding creating solutions prematurely. Kent uses examples like Tesla's electric engine and Remix framework to illustrate the benefits of problem elimination. He emphasizes the value of trade-offs and taking the easier path, as well as the need to constantly re-evaluate and change approaches to eliminate problems.
Using useEffect Effectively
React Advanced 2022React Advanced 2022
30 min
Using useEffect Effectively
Top Content
Today's Talk explores the use of the useEffect hook in React development, covering topics such as fetching data, handling race conditions and cleanup, and optimizing performance. It also discusses the correct use of useEffect in React 18, the distinction between Activity Effects and Action Effects, and the potential misuse of useEffect. The Talk highlights the benefits of using useQuery or SWR for data fetching, the problems with using useEffect for initializing global singletons, and the use of state machines for handling effects. The speaker also recommends exploring the beta React docs and using tools like the stately.ai editor for visualizing state machines.
Design Systems: Walking the Line Between Flexibility and Consistency
React Advanced 2021React Advanced 2021
47 min
Design Systems: Walking the Line Between Flexibility and Consistency
Top Content
The Talk discusses the balance between flexibility and consistency in design systems. It explores the API design of the ActionList component and the customization options it offers. The use of component-based APIs and composability is emphasized for flexibility and customization. The Talk also touches on the ActionMenu component and the concept of building for people. The Q&A session covers topics such as component inclusion in design systems, API complexity, and the decision between creating a custom design system or using a component library.
React Concurrency, Explained
React Summit 2023React Summit 2023
23 min
React Concurrency, Explained
Top Content
React 18's concurrent rendering, specifically the useTransition hook, optimizes app performance by allowing non-urgent updates to be processed without freezing the UI. However, there are drawbacks such as longer processing time for non-urgent updates and increased CPU usage. The useTransition hook works similarly to throttling or bouncing, making it useful for addressing performance issues caused by multiple small components. Libraries like React Query may require the use of alternative APIs to handle urgent and non-urgent updates effectively.
Managing React State: 10 Years of Lessons Learned
React Day Berlin 2023React Day Berlin 2023
16 min
Managing React State: 10 Years of Lessons Learned
Top Content
This Talk focuses on effective React state management and lessons learned over the past 10 years. Key points include separating related state, utilizing UseReducer for protecting state and updating multiple pieces of state simultaneously, avoiding unnecessary state syncing with useEffect, using abstractions like React Query or SWR for fetching data, simplifying state management with custom hooks, and leveraging refs and third-party libraries for managing state. Additional resources and services are also provided for further learning and support.
TypeScript and React: Secrets of a Happy Marriage
React Advanced 2022React Advanced 2022
21 min
TypeScript and React: Secrets of a Happy Marriage
Top Content
React and TypeScript have a strong relationship, with TypeScript offering benefits like better type checking and contract enforcement. Failing early and failing hard is important in software development to catch errors and debug effectively. TypeScript provides early detection of errors and ensures data accuracy in components and hooks. It offers superior type safety but can become complex as the codebase grows. Using union types in props can resolve errors and address dependencies. Dynamic communication and type contracts can be achieved through generics. Understanding React's built-in types and hooks like useState and useRef is crucial for leveraging their functionality.

Workshops on related topic

React Performance Debugging Masterclass
React Summit 2023React Summit 2023
170 min
React Performance Debugging Masterclass
Top Content
Featured Workshop
Ivan Akulov
Ivan Akulov
Ivan’s first attempts at performance debugging were chaotic. He would see a slow interaction, try a random optimization, see that it didn't help, and keep trying other optimizations until he found the right one (or gave up).
Back then, Ivan didn’t know how to use performance devtools well. He would do a recording in Chrome DevTools or React Profiler, poke around it, try clicking random things, and then close it in frustration a few minutes later. Now, Ivan knows exactly where and what to look for. And in this workshop, Ivan will teach you that too.
Here’s how this is going to work. We’ll take a slow app → debug it (using tools like Chrome DevTools, React Profiler, and why-did-you-render) → pinpoint the bottleneck → and then repeat, several times more. We won’t talk about the solutions (in 90% of the cases, it’s just the ol’ regular useMemo() or memo()). But we’ll talk about everything that comes before – and learn how to analyze any React performance problem, step by step.
(Note: This workshop is best suited for engineers who are already familiar with how useMemo() and memo() work – but want to get better at using the performance tools around React. Also, we’ll be covering interaction performance, not load speed, so you won’t hear a word about Lighthouse 🤐)
React Hooks Tips Only the Pros Know
React Summit Remote Edition 2021React Summit Remote Edition 2021
177 min
React Hooks Tips Only the Pros Know
Top Content
Featured Workshop
Maurice de Beijer
Maurice de Beijer
The addition of the hooks API to React was quite a major change. Before hooks most components had to be class based. Now, with hooks, these are often much simpler functional components. Hooks can be really simple to use. Almost deceptively simple. Because there are still plenty of ways you can mess up with hooks. And it often turns out there are many ways where you can improve your components a better understanding of how each React hook can be used.You will learn all about the pros and cons of the various hooks. You will learn when to use useState() versus useReducer(). We will look at using useContext() efficiently. You will see when to use useLayoutEffect() and when useEffect() is better.
React, TypeScript, and TDD
React Advanced 2021React Advanced 2021
174 min
React, TypeScript, and TDD
Top Content
Featured Workshop
Paul Everitt
Paul Everitt
ReactJS is wildly popular and thus wildly supported. TypeScript is increasingly popular, and thus increasingly supported.

The two together? Not as much. Given that they both change quickly, it's hard to find accurate learning materials.

React+TypeScript, with JetBrains IDEs? That three-part combination is the topic of this series. We'll show a little about a lot. Meaning, the key steps to getting productive, in the IDE, for React projects using TypeScript. Along the way we'll show test-driven development and emphasize tips-and-tricks in the IDE.
Master JavaScript Patterns
JSNation 2024JSNation 2024
145 min
Master JavaScript Patterns
Top Content
Featured Workshop
Adrian Hajdin
Adrian Hajdin
During this workshop, participants will review the essential JavaScript patterns that every developer should know. Through hands-on exercises, real-world examples, and interactive discussions, attendees will deepen their understanding of best practices for organizing code, solving common challenges, and designing scalable architectures. By the end of the workshop, participants will gain newfound confidence in their ability to write high-quality JavaScript code that stands the test of time.
Points Covered:
1. Introduction to JavaScript Patterns2. Foundational Patterns3. Object Creation Patterns4. Behavioral Patterns5. Architectural Patterns6. Hands-On Exercises and Case Studies
How It Will Help Developers:
- Gain a deep understanding of JavaScript patterns and their applications in real-world scenarios- Learn best practices for organizing code, solving common challenges, and designing scalable architectures- Enhance problem-solving skills and code readability- Improve collaboration and communication within development teams- Accelerate career growth and opportunities for advancement in the software industry
Designing Effective Tests With React Testing Library
React Summit 2023React Summit 2023
151 min
Designing Effective Tests With React Testing Library
Top Content
Featured Workshop
Josh Justice
Josh Justice
React Testing Library is a great framework for React component tests because there are a lot of questions it answers for you, so you don’t need to worry about those questions. But that doesn’t mean testing is easy. There are still a lot of questions you have to figure out for yourself: How many component tests should you write vs end-to-end tests or lower-level unit tests? How can you test a certain line of code that is tricky to test? And what in the world are you supposed to do about that persistent act() warning?
In this three-hour workshop we’ll introduce React Testing Library along with a mental model for how to think about designing your component tests. This mental model will help you see how to test each bit of logic, whether or not to mock dependencies, and will help improve the design of your components. You’ll walk away with the tools, techniques, and principles you need to implement low-cost, high-value component tests.
Table of contents- The different kinds of React application tests, and where component tests fit in- A mental model for thinking about the inputs and outputs of the components you test- Options for selecting DOM elements to verify and interact with them- The value of mocks and why they shouldn’t be avoided- The challenges with asynchrony in RTL tests and how to handle them
Prerequisites- Familiarity with building applications with React- Basic experience writing automated tests with Jest or another unit testing framework- You do not need any experience with React Testing Library- Machine setup: Node LTS, Yarn
Next.js 13: Data Fetching Strategies
React Day Berlin 2022React Day Berlin 2022
53 min
Next.js 13: Data Fetching Strategies
Top Content
Workshop
Alice De Mauro
Alice De Mauro
- Introduction- Prerequisites for the workshop- Fetching strategies: fundamentals- Fetching strategies – hands-on: fetch API, cache (static VS dynamic), revalidate, suspense (parallel data fetching)- Test your build and serve it on Vercel- Future: Server components VS Client components- Workshop easter egg (unrelated to the topic, calling out accessibility)- Wrapping up