TL;DR

Google Jules is the only major coding agent built around queuing instead of live chat. You describe a task, walk away, and a pull request shows up later. The free tier gives 15 tasks per day on Gemini 3 Flash. The $19.99/month Pro tier bumps that to 100 tasks on Gemini 3.1 Pro, and proactive features like CI Fixer and Scheduled Tasks make it feel less like a tool and more like a junior developer who never goes offline. But Jules is slow, can’t handle files over ~50K lines, and only connects to GitHub. If you need real-time pair programming or work with GitLab, look elsewhere.

Why an Async Coding Agent

Real-time terminal agents like Claude Code and Codex CLI work the way the name suggests: you type a prompt and watch code materialize. They’re good at that. The friction shows up when you have several independent tasks. You end up babysitting the agent through each one in sequence, switching between planning the work and watching it type.

Jules promises something different. Describe the task, hit submit, go do something else, and come back to a pull request. The Pro tier costs $19.99 a month, bundled with Google AI Pro, and it’s pitched at backlog work: dependency bumps, test scaffolding, small bug fixes.

The short version: Jules delivered on the async promise. But “async” also means “slow,” and the tradeoffs stack up in ways the marketing doesn’t mention.

How Jules Works

Every task runs in an isolated Google Cloud VM. Jules clones your repo, reads the codebase, builds an execution plan, and shows you that plan before touching any files. You can edit the plan, approve it, or scrap it entirely. Once approved, Jules works through the changes file by file, running any tests it finds at each step. When it’s done, it opens a PR on GitHub.

The whole loop is: submit, approve a plan, wait for the PR notification. No terminal session, no streaming output, no watching characters appear.

15–300
Tasks/day by tier
3–60
Concurrent tasks
80.6%
SWE-bench (Gemini 3.1 Pro)

The model underneath depends on your tier. Free gets Gemini 3 Flash. Pro and Ultra run Gemini 3.1 Pro, which scores 80.6% on SWE-bench Verified — competitive with Claude Opus 4.6 at 80.8%, though behind the current flagship Opus 4.8’s 88.6% in agentic scaffolding. (For a full breakdown of how these models compare on coding tasks, see the GPT-5.4 vs Claude Opus 4.7 vs Gemini 3.1 Pro comparison.)

Pricing Breakdown

Jules doesn’t have its own subscription. It bundles into Google’s AI tiers:

FreePro ($19.99/mo)Ultra ($99.99/mo)
Daily tasks15100300
Concurrent tasks31560
ModelGemini 3 FlashGemini 3.1 ProGemini 3.1 Pro (priority)
Suggested TasksNoYesYes
Scheduled TasksYesYesYes

The free tier is generous enough for evaluation. Fifteen tasks per day covers most solo developers who want to offload grunt work. Pro makes sense once you’re running 10+ tasks daily and want the model upgrade. Ultra is for teams running agent-heavy workflows — 60 concurrent tasks means you can point Jules at an entire sprint backlog and let it churn.

One catch: paid plans require a @gmail.com account. Google Workspace users can’t subscribe yet.

What Jules Got Right

Batch Parallelism

The async model isn’t just a UX gimmick. Queue five dependency-bump tasks in the morning, go write the design doc you’ve been putting off, and come back to five PRs. In a real-time agent, the same five tasks mean approving file edits and answering clarification prompts one at a time.

On Pro, 15 concurrent slots mean a whole backlog can run in parallel without queueing. That’s the real difference from a terminal agent that works through one task at a time.

CI Fixer

This is the standout feature. When a GitHub Actions workflow fails, Jules automatically analyzes the logs, writes a fix, commits it, and resubmits. It loops until CI passes or gives up after a configurable number of attempts.

CI failures are the clearest use case. Point Jules at a failing build, say tests broken by a deprecated API after a library upgrade, and it reads the logs, traces the failing call, patches it and pushes a new build for review. It fits the fixes that are tedious rather than hard.

Scheduled Tasks

You can set Jules to run recurring jobs: nightly lint passes, weekly dependency audits, monthly dead-code sweeps. This is the part that makes Jules feel like a team member rather than a tool. A weekly pip-audit job, for example, turns a check most teams run once a quarter into a Monday-morning PR with any new CVEs patched.

Suggested Tasks

On Pro and Ultra, Jules scans up to five repos and proposes improvements. Stale TODO comments are an easy first target: it finds forgotten # TODO: handle edge case annotations and opens PRs that actually handle them.

The suggestions aren’t always useful. Jules proposed refactoring a perfectly fine utility function into a class hierarchy that added complexity for zero benefit. But the hit rate was around 60-70%, and dismissing bad suggestions takes seconds.

Where Jules Falls Short

Speed

Jules is slow. A task that Claude Code handles in 90 seconds takes Jules 8-15 minutes. Part of this is the VM spin-up, part is the planning phase (Jules builds a detailed plan before writing any code), and part is that Gemini 3.1 Pro generates tokens slower than Claude in agentic loops.

For anything urgent (a production bug, a quick fix before a demo) Jules isn’t the right tool. You’ll be staring at a progress bar while Claude Code would have already pushed the commit.

Large File Blindness

Gemini 3.1 Pro has a 1M-token context window, but Jules appears to work with a tighter limit in practice, and very large files are where that shows. On the multi-thousand-line handler modules legacy services accumulate, its plans can reference functions it never saw, a sign it’s working from a truncated view.

Real-time agents handle this differently. Claude Code can stream file reads and focus on specific sections. Jules loads the whole context upfront and chokes on anything too large.

GitHub Only

No GitLab. No Bitbucket. No self-hosted Git. If your repos aren’t on github.com, Jules can’t touch them. Google Workspace integration is also missing, which means enterprise teams on Google Cloud who use Cloud Source Repositores are locked out too.

Language Coverage

Python and TypeScript/JavaScript are first-class citizens. Jules writes solid code in both, catches edge cases, and uses idiomatic patterns. Go, Java, and C# work but with noticeably lower reliability. My Go microservices got PRs that compiled but missed patterns any Go developer would catch: unchecked errors, bare returns where wrapped errors belong.

Hallucinated Progress

Check that a “completed” task really completed. Async agents can report success on a run that stalled partway, leaving a PR with half the files edited and the tests never run, and the UI doesn’t make that obvious. You find out in code review, which undercuts the “queue and forget” promise. If you’re relying on any coding agent for unsupervised work, setting up guardrails before you go hands-off is worth the time.

Jules vs the Competition

FeatureGoogle JulesClaude CodeGitHub Copilot AgentOpenAI Codex
Interaction modelAsync (queue + PR)Real-time terminalBoth (IDE + async)Async (cloud tasks)
Pricing$0–99.99/mo$20/mo (Pro) or API$10–39/moAPI-based
ModelGemini 3.1 ProClaude Opus 4.8GPT-5.6-Codex (default)GPT-5.6-Codex
SWE-bench80.6%88.6%~77–80%85%
Concurrent tasks3–601 (serial)1–3Varies
Proactive featuresCI Fixer, ScheduledNoneLimitedNone
Git platformsGitHub onlyAnyGitHub onlyGitHub only
Best forBatch work, maintenanceComplex refactors, explorationGitHub-native workflowsAutomated fixes

What counts here is workflow fit, not a feature checklist.

Jules owns the batch maintenance lane. Queue 20 dependency bumps and lint fixes, check the PRs over coffee. On Pro with 15 concurrent slots, a full day’s grunt work finishes before lunch. No other agent handles this volume as smoothly.

Claude Code is the better pick for anything that needs back-and-forth. Debugging a race condition, designing an API, exploring unfamiliar code — you want a real-time thinking partner, and Opus 4.8’s 8-point SWE-bench lead over Gemini 3.1 Pro shows up when the task gets hard. (I covered the DeepSeek V4 Pro review recently, and it’s another strong option at a fraction of Claude’s API cost.)

Copilot Agent fits if you already live in GitHub Issues and Actions. It’s the least friction for teams whose entire workflow is PR-centric.

Where Jules pulls ahead of all three: proactive features. Among the agents compared here, only Jules offers CI auto-fixing and scheduled recurring tasks, and for a team with a steady maintenance backlog that alone can justify the Pro tier.

MCP Server Integration

In February 2026, Jules added Model Context Protocol support with six hand-selected servers: Linear, Stitch, Neon, Tinybird, Context7, and Supabase. Google took a curated approach: every server was audited for data flow and tool permissions before being allowed.

In practice, this means Jules can read your Linear tickets, query your Neon database schema, and check Supabase auth configuration while planning changes. With the Neon MCP server connected, a task like “add pagination to the /users endpoint based on the current schema” starts from the real schema instead of one you paste into the task description.

Six servers is limiting. Claude Code connects to any MCP server you configure. But Google’s curated approach makes sense for an agent that runs in a cloud VM with repo access. A malicious MCP server could exfiltrate code, so restriction buys you something real.

The Jules API

Google also launched a Jules API for programmatic task creation. You can trigger Jules tasks from CI pipelines, chatbots, or custom tooling. The API exposes task creation, status polling, and result retrieval.

As of this review the API is in v1alpha, so field names and auth methods may change. Here’s the general shape of a session-creation call using the current schema:

import requests

API_KEY = "your-google-api-key"

session = requests.post(
    "https://jules.googleapis.com/v1alpha/sessions",
    headers={"X-Goog-Api-Key": API_KEY},
    json={
        "sourceContext": {
            "gitHub": {"repository": "owner/repo", "branch": "main"}
        },
        "title": "Add input validation to /users POST endpoint",
    },
)
print(session.json())
# {"name": "sessions/abc123", "state": "CREATED", ...}

The automationMode field controls whether Jules runs without human review of its execution plan. The default, manual approval, is the safer place to start: you see the plan before Jules edits any files. For trusted, repeatable tasks like dependency bumps, switching to full automation turns Jules into an autonomous pipeline.

The obvious next step is connecting Jules to your issue tracker: new bug filed, Jules automatically attempts a fix, PR shows up for review. The Stitch design team at Google reportedly runs “a pod of daily Jules agents” with assigned roles (performance tuning, security patching, accessibility, test coverage), making Jules, according to the team’s blog post, one of the largest contributors to their repository. That kind of multi-agent delegation is still rare — Anthropic’s 2026 agentic coding report found a stark gap between AI usage (60%) and actual delegation (0-20%), suggesting most teams aren’t yet trusting agents to run unsupervised the way the Stitch team does.

Project Jitro: The Goal-Driven Successor

Google previewed Project Jitro at I/O 2026 — the next version of Jules that shifts from task-driven to goal-driven. Instead of “fix this function,” you’d say “get test coverage to 85%” and Jitro figures out which files to change, which tests to write, and how to get the metric where you want it.

The current Jules already hints at this direction. Suggested Tasks, Scheduled Tasks, and the Render integration all share one pattern: Jules initiating action based on codebase state. Jitro takes that to its logical conclusion.

The obvious question is accountability. When an agent autonomously refactors modules to hit a metric, who reviews the architectural decisions it made along the way? Google hasn’t answered that yet. Jitro was previewed under a waitlist at I/O 2026, and Google hadn’t announced a general-availability date as of this review.

Who Should Use Jules

Good fit:

  • You maintain multiple repos and spend hours weekly on dependency updates, lint fixes, and test scaffolding
  • You want CI failures fixed automatically without context-switching from whatever you’re building
  • You work in Python or TypeScript and your repos are on GitHub
  • You like reviewing PRs more than supervising an agent in real time

Skip it:

  • You need real-time collaboration — architecture discussions, exploratory coding, debugging complex state
  • Your repos are on GitLab, Bitbucket, or self-hosted Git
  • You work primarily in Go, Java, or C# where Jules’s output needs heavy review anyway
  • You need to work with files over 50K lines

FAQ

Is Google Jules free?

Yes, the free tier gives 15 tasks per day with 3 concurrent slots, running on Gemini 3 Flash. No credit card required. It’s enough to evaluate whether the async model fits your workflow before committing to Pro.

How does Google Jules compare to Claude Code?

They solve different problems. Jules is async — you queue tasks and get PRs back later. Claude Code is real-time — you work together in a terminal session. Jules is better for batch maintenance work across multiple repos. Claude Code is better for complex single-task work where you need back-and-forth. Claude’s underlying model (Opus 4.8, 88.6% SWE-bench) also outperforms Jules’s Gemini 3.1 Pro (80.6%) on coding benchmarks.

What languages does Google Jules support?

Python and TypeScript/JavaScript are best supported. Go, Java, and C# work but produce less reliable output. Expect to catch missed error handling patterns and non-idiomatic code during review.

Can Jules work with private repositories?

Yes. Jules clones repos into isolated Google Cloud VMs. Google states your code isn’t used for model training. The VM is ephemeral — spun up per task and destroyed after.

What is Project Jitro?

Project Jitro is Google’s next-generation coding agent, previewed at I/O 2026. Instead of describing a task (“fix this bug”), you define a goal (“reduce p95 latency by 30ms”) and the agent determines the changes needed. It was on a waitlist at I/O 2026, with no general-availability date announced as of this review.

Sources

Bottom Line

Jules is the best coding agent for people who hate babysitting coding agents. The async model, CI Fixer, and Scheduled Tasks create a workflow where maintenance work runs on autopilot. Set up well, Monday mornings start with a few PRs from overnight audit and lint runs. For $19.99/month, that trade works.

For thinking-partner work (debugging a race condition, designing an API, exploring unfamiliar code) you still need Claude Code or Copilot. Jules takes orders and delivers results, on its own schedule, at its own pace.

If your bottleneck is “too many small tasks, not enough hands,” try the free tier for a week. Queue up your backlog. See what comes back. The 15-task daily limit is enough to know whether this fits your workflow.