Google Jules vs Codex: Which Async Coding Agent Wins in 2026?
| Tool | Rating | Price | Best For | Action |
|---|---|---|---|---|
GJ Google Jules | 4.3 | Free (15 tasks/day) / $19.99/mo via Google AI Pro / $124.99/mo via Google AI Ultra | Try Google Jules Free | |
C Codex | 4.7 | Free / $8/mo Go / $20/mo Plus / $100–$500/mo Pro / $20/user/mo Business | Try Codex Free |
Google Jules vs Codex: Which Async Coding Agent Wins in 2026?
Google Jules and OpenAI Codex are the two most prominent asynchronous coding agents of 2026. Neither is a tab-completion tool or an in-editor chat sidebar. You hand both of them a task, walk away, and come back to a pull request.
But they got to that shared destination from opposite directions. Jules is a GitHub-native agent from Google Labs that lives almost entirely in the cloud. Codex is a five-surface product — web, CLI, IDE extension, iOS and cloud — that treats async delegation as one mode among several.
The short version: Jules has the better free tier and the stronger CI/CD automation story. Codex has the better interactive control, wider integration surface, and a model lineup that lets you tune cost against capability. If you are a solo developer on a Gmail account who wants background PRs for free, start with Jules. If you want one agent that handles both delegated and hands-on work, Codex is the safer bet.
Quick Comparison
| Feature | Google Jules | Codex |
|---|---|---|
| Entry price | Free (15 tasks/day) | Free (minimal) / $8/mo Go |
| Mainstream tier | $19.99/mo (Google AI Pro) | $20/mo (ChatGPT Plus) |
| Top tier | $124.99/mo (Google AI Ultra) | $500/mo Pro |
| Team plan | Not available on Workspace accounts | $20/user/mo Business (annual, 2+ seats) |
| Models | Gemini 2.5 Pro (free), Gemini 3.1 Pro (paid) | GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna |
| Surfaces | Web, Jules Tools CLI, public API | Web, CLI, IDE extension, iOS, cloud |
| Concurrency | 3 / 15 / 60 tasks by tier | 5-hour rolling windows + weekly caps |
| Steer mid-task | No | Yes (instant_interrupt in CLI) |
| Task launchers | GitHub, CLI, API | GitHub, GitLab, Linear, Slack, web |
| MCP support | Six vetted servers | Yes, incl. OAuth-secret servers |
| Auto code review | CI Fixer for failed Actions | Automatic first pass on PRs/MRs |
| Memory | Per-repository agentic memory | Reusable dev environments |
What Each Tool Actually Is
Jules: Google's GitHub-Native Async Agent
Jules describes itself as "an experimental coding agent that helps you fix bugs, add documentation, and build new features." The workflow is deliberately hands-off:
- You connect a GitHub repository and assign a task.
- Jules clones the code into a cloud virtual machine and installs dependencies.
- It produces a plan and waits for your approval before touching any code.
- It executes, runs tests, and opens a pull request.
- A browser notification tells you it's done — or that it needs input.
Jules reads AGENTS.md files to understand repository conventions, which is the same convention Codex and most 2026 agents now follow.
Two 2026 additions changed how Jules fits into real teams. In February 2026 it gained MCP server support — initially Linear, Stitch, Neon, Tinybird, Context7 and Supabase — plus a CI Fixer that detects failed GitHub Actions checks and pushes fixes without being asked, and account-level commit authorship control (Jules only, co-authored, or you only). In March 2026, Gemini 3.1 Pro became the default model for Google AI Pro subscribers.
Google's Gemini 3 rollout also brought three capabilities worth naming:
- Coherent planning — multi-step tasks hold together across transitions with less babysitting.
- Visual verification — Jules renders and checks web app output using Gemini's multimodal capability rather than trusting the diff.
- Agentic memories — Jules saves your corrections and nudges per repository and applies them to similar future tasks.
Outside the web app, Jules Tools is a CLI that lets you start, stop and check tasks alongside your own terminal commands, and the Jules API (early preview) lets you drop Jules into GitHub Actions pipelines or wire it into Slack, Linear and Jira.
Codex: OpenAI's Multi-Surface Coding Agent
Codex is broader in scope. It runs in the ChatGPT app, an IDE extension, a CLI, on iOS, and in Codex cloud — and the cloud version accepts tasks launched from the web, GitHub, GitLab, Linear or Slack. Teams can define reusable development environments so everyone's cloud tasks start from the same approved setup and permissions.
Codex's 2026 model lineup is three-tiered rather than one-size-fits-all:
- GPT-6 Astra — the heavyweight, most expensive per token.
- GPT-6.1 Sol — launched at DevDay on 29 September 2026; near-Astra performance for complex work at lower cost, explicitly aimed at repeated long-running work.
- GPT-6 Luna — cheap and fast for high-volume, simpler work.
The Code Review experience in the ChatGPT desktop app takes an automatic first pass on changes in the cloud and posts feedback directly to GitHub pull requests or GitLab merge requests.
And critically for the async-vs-interactive question: Codex CLI 0.159.0 (September 2026) shipped opt-in instant_interrupt, which lets you steer Codex during a model response. The October 2026 release added an agent command center for browsing older tasks and opt-in Guardian review.
Pricing: Bundled Subscription vs Credit Tiers
Jules Pricing (verified July–August 2026)
| Tier | Cost | Daily tasks | Concurrent | Model |
|---|---|---|---|---|
| Free | $0 | 15 | 3 | Gemini 2.5 Pro |
| Jules in Pro | $19.99/mo (Google AI Pro) | 100 | 15 | Gemini 3.1 Pro |
| Jules in Ultra | $124.99/mo (Google AI Ultra) | 300 | 60 | Gemini 3 Pro, priority access |
Jules is not sold standalone. The paid tiers come bundled with Google AI Pro and Google AI Ultra subscriptions, which also include Gemini app access, storage and other Google AI features — so the effective cost depends on whether you'd buy those anyway.
There is one restriction that disqualifies Jules outright for many teams: paid Jules plans are only available on individual Google accounts ending in @gmail.com. A company Google Workspace account cannot upgrade. If your organisation runs on Workspace, your developers are capped at the free tier.
Codex Pricing
| Plan | Cost | Notes |
|---|---|---|
| Free | $0 | Basic capabilities only |
| Go | $8/mo | Lightweight coding tasks |
| Plus | $20/mo | Codex on web, CLI, IDE and iOS; GPT-6.1 Sol + GPT-6 Luna |
| Pro | $100 / $200 / $500/mo | $500 tier adds Astra Ultrafast mode |
| Business | $20/user/mo (annual, 2+ users) | Admin controls, larger cloud VMs |
| Enterprise / Edu | Custom | Full org deployment, larger VMs |
| API | Pay-per-use | No cloud features |
Codex meters usage on a rolling five-hour window plus a weekly cap. Estimated local messages per five-hour period: GPT-6 Astra 5–45, GPT-6.1 Sol 15–160, GPT-6 Luna 350–3,000.
Credit rates per million tokens at Standard speed:
| Model | Input | Output |
|---|---|---|
| GPT-6 Astra | 250 | 1,250 |
| GPT-6.1 Sol | 50 | 250 |
| GPT-6 Luna | 2.5 | 12.5 |
Speed modes multiply consumption — Fast is 2x, and Ultrafast is 6x for Astra. That 100x spread between Astra and Luna is Codex's real pricing story: model choice matters more than plan choice. Via the API, GPT-6.1 Sol lists at $2 per million input tokens and $10 per million output.
Cost Verdict
For light use, Jules wins decisively — 15 real tasks per day for $0 has no Codex equivalent. At the mainstream tier the two are effectively tied on headline price ($19.99 vs $20), but they meter differently: Jules counts whole tasks, Codex counts tokens weighted by model. Jules' limits are far easier to predict. Codex's are far easier to optimise.
For teams, Codex wins by default, because Jules has no Workspace-compatible paid plan. $20/user/month Business is the only one of the two that an IT department can actually buy.
Models and Benchmarks
Gemini 3.1 Pro — the model behind paid Jules — scores 80.6% on SWE-bench Verified, which puts it essentially level with Claude Opus 4.6 at 80.8%. That is frontier-class performance on real GitHub issue resolution, and it is the strongest argument for Jules on pure capability.
On the OpenAI side, the picture is more about cost-efficiency than a single headline number. OpenAI reports that GPT-6.1 Sol more than doubles GPT-6 Sol's score on Terminal-Bench Science 0.1 at maximum reasoning effort, averaging $5.47 per task. On independent Terminal-Bench 4.0, Claude Opus 5.5 holds the top score, with GPT-6 Astra close behind at under half the cost per task.
Read benchmarks carefully here. SWE-bench Verified measures the model; Terminal-Bench measures the agent harness around it. A strong model inside a slow harness still produces a slow agent — which brings us to the single biggest practical difference between these two tools.
Speed and Control: The Real Dividing Line
This is where the comparison stops being close.
Jules is slow. Reviewers consistently report 8–15 minutes for tasks that terminal-native agents finish in around 90 seconds. Three things stack up: VM spin-up time, the mandatory planning phase before any code is written, and Gemini 3.1 Pro generating tokens more slowly than competitors inside agentic loops. For a production incident or a fix before a demo, Jules is the wrong tool.
Jules also cannot be redirected. Once a task starts, you find out it went the wrong way when the PR arrives. That is the defining constraint of its async model, and it compounds a second weakness: Jules struggles with ambiguous issues. "The login is broken" makes Jules guess. A concrete bug with reproduction steps, or a feature with a stated interface, works far better. It validates its work against your tests — so thin test coverage means you are reviewing unknown unknowns.
Codex is not immune to any of this when running cloud tasks. But it has an escape hatch Jules doesn't: instant_interrupt in the CLI lets you steer mid-response, and the IDE and CLI surfaces let you work interactively when a task turns out to need a human. You can downgrade a Codex task from async to hands-on. You cannot do that with Jules.
One more asymmetry: Gemini 3.1 Pro advertises a 1M-token context window, but reviewers report Jules imposes a tighter effective limit in practice, with very large files off-limits. Big monorepos expose this quickly.
Integrations and Ecosystem
Jules connects through GitHub, the Jules Tools CLI and the public API. The API is the interesting piece — it's built for CI/CD, so you can trigger Jules from GitHub Actions or embed it into Slack, Linear, Jira and GitHub. The CI Fixer closes that loop: a failed check becomes a fix without a human in the middle. MCP support exists but is restricted to six hand-selected servers (Linear, Stitch, Neon, Tinybird, Context7, Supabase), authenticated with API keys in Settings. Google frames this as security-first; in practice it limits what you can wire up compared to open-MCP agents.
Codex accepts cloud task launches from GitHub, GitLab, Linear, Slack and the web, and its MCP support extends to servers requiring pre-registered OAuth client secrets — a meaningful enterprise detail. Recent CLI releases added bearer-token support for secure WebSocket connections and terminal input approval on by default.
Verdict: Jules wins on autonomous CI/CD — the CI Fixer has no direct Codex equivalent. Codex wins on breadth, GitLab support, and MCP flexibility.
Who Should Use Each Tool
Choose Jules if you:
- Want a substantial free tier — 15 tasks/day with 3 concurrent is real capacity
- Have a personal Gmail account, not a Workspace one
- Already pay for Google AI Pro or Ultra and want the agent included
- Want massive concurrency — 60 parallel tasks on Ultra is unmatched at the price
- Want CI failures fixed automatically without wiring anything yourself
- Work on well-specified issues with solid test coverage
- Prefer reviewing a plan before code gets written
Choose Codex if you:
- Need both delegated async work and hands-on interactive coding
- Want to steer an agent mid-task instead of waiting for a wrong answer
- Work on GitLab, or launch tasks from Slack or Linear
- Run a team on Google Workspace — Jules paid plans are closed to you
- Want to tune cost per task via Luna/Sol/Astra model choice
- Need automatic code review posted to PRs and MRs
- Want mobile access to running tasks via iOS
The Hybrid Case
These two genuinely stack. Jules' free tier costs nothing, so there's no reason not to point it at your low-stakes backlog — dependency bumps, documentation, flaky CI checks — while Codex handles the work you need to supervise. The only real blocker is the Workspace restriction, which forces Workspace-based teams into Jules' free tier whether they like it or not.
Final Verdict
Jules is the more opinionated product, and the opinion is "trust the agent." The free tier is the most generous in this category, Gemini 3.1 Pro is frontier-class on SWE-bench Verified, and the CI Fixer plus public API make it the better pure-automation tool. But the 8–15 minute task times, the inability to course-correct, the practical context limits, and above all the @gmail.com-only paid tiers cap how far it can go in a real engineering organisation.
Codex is the more complete tool. Five surfaces, mid-response steering, GitLab and Slack launchers, automatic code review, a buyable team plan, and a three-model lineup that lets you spend $2.50 or $250 per million input tokens depending on the job. The cost is complexity — credit-based billing with 2x and 6x speed multipliers is genuinely hard to forecast, and you're locked to GPT models.
Our recommendation: Try Jules first, because it's free and the evaluation costs you nothing but a GitHub connection. Use it for well-scoped, non-urgent work. But if you're buying one async coding agent for a team — especially a team on Google Workspace — buy Codex. The ability to take over a task mid-flight is worth more in practice than any benchmark gap between Gemini 3.1 Pro and GPT-6.1 Sol.
Pricing, limits and model availability accurate as of October 2026. AI coding agents change monthly — check the official pricing pages before committing.
Sources: Jules docs, Jules changelog, Jules pricing breakdown (Hackup), Building with Gemini 3 in Jules, VentureBeat on Jules CLI and API, ChatGPT & Codex pricing, ChatGPT & Codex changelog, GPT-6.1 Sol at DevDay (Unite.AI), Codex Cloud explained (Agent37), Gemini 3.1 Pro benchmarks (BenchLM), Terminal-Bench agent rankings (Morph)
Pros
- Genuinely usable free tier — 15 tasks/day, 3 concurrent
- Up to 60 concurrent tasks on Ultra
- CI Fixer auto-repairs failed GitHub Actions checks
- Plan-review step before any code is written
- Agentic memory learns per-repository preferences
- Jules Tools CLI and public API for CI/CD pipelines
Cons
- Slow — 8–15 minutes per task is normal
- Cannot be redirected once a task starts
- Paid tiers require a personal @gmail.com account — Workspace blocked
- MCP support limited to six vetted servers
- Struggles with ambiguous issues and very large files
Pros
- Five surfaces: web, CLI, IDE extension, iOS, cloud
- instant_interrupt lets you steer mid-response in the CLI
- Cloud tasks launch from GitHub, GitLab, Linear or Slack
- Automatic code-review pass on PRs and MRs
- Three-model lineup lets you trade cost against capability
- Reusable team dev environments with approved permissions
Cons
- No free-tier substance — real work needs $20/mo+
- Credit-based billing is hard to forecast
- GPT models only — no Gemini or Claude
- Speed modes multiply credit burn up to 6x
- API access excludes all the cloud features