Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Run Gemma 4 31B for Image, Video and Reasoning Tasks

Google DeepMind's largest dense Gemma 4 model handles images, video frames and 256K-token context. Here's how to run it with Transformers.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Gemma 4 31B-it is the instruction-tuned version of the largest dense model in Google DeepMind’s Gemma 4 family. It takes text, images and video (as frames) as input and returns text. It has a 256K-token context window and a thinking mode you can turn on or off. It is released under Apache 2.0. The main reason to look at it is how much better it scores than the previous generation. On the card’s benchmark table it scores 80.0% on LiveCodeBench v6 and 89.2% on AIME 2026 without tools. Gemma 3 27B, run without thinking, scored 29.1% and 20.8% on those benchmarks. The card doesn’t say whether thinking was on for the Gemma 4 scores, so some of that gap may come from thinking rather than from the new generation.

Key specs

Property Gemma 4 31B
Parameters 30.7B (dense)
Layers 60
Sliding window 1024 tokens
Context length 256K tokens
Vocabulary 262K
Modalities Text, image (video as frames)
Vision encoder ~550M parameters
Audio Not supported
Training data cutoff January 2025

Selected instruction-tuned results from the card:

Benchmark 31B 26B A4B (MoE) Gemma 3 27B (no think)
MMLU Pro 85.2% 82.6% 67.6%
LiveCodeBench v6 80.0% 77.1% 29.1%
Codeforces ELO 2150 1718 110
GPQA Diamond 84.3% 82.3% 42.4%
Tau2 (avg over 3) 76.9% 68.2% 16.2%
MMMU Pro 76.9% 73.8% 49.7%
OmniDocBench 1.5 (edit distance, lower is better) 0.131 0.149 0.365
MRCR v2 8 needle 128k 66.4% 44.1% 13.5%

The model uses a hybrid attention design. Local sliding-window layers are interleaved with global layers, and the final layer is always global. The global layers use unified keys and values with Proportional RoPE to save memory on long contexts.

Install

Terminal window
pip install -U transformers torch torchvision accelerate

The card lists torchvision for image input and adds librosa for its video example:

Terminal window
pip install -U transformers torch torchvision librosa accelerate

Run it

chat.py
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "google/gemma-4-31B-it"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Write a short joke about saving RAM."},
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
enable_thinking=False
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
outputs = model.generate(**inputs, max_new_tokens=1024)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(processor.parse_response(response, prefix=inputs["input_ids"]))

The system role is supported natively in Gemma 4. Gemma 3 did not have it.

Reasoning with thinking mode

For math, coding and multi-step problems, set enable_thinking=True. The card says parse_response handles separating the thinking output from the final answer. The card doesn’t say which thinking setting was used for its benchmark scores, so treat this as a sensible default for hard problems, not as the configuration behind the numbers above.

reason.py
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "google/gemma-4-31B-it"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)
problem = (
"A function f satisfies f(x + y) = f(x) + f(y) + 2xy for all integers x, y, "
"and f(1) = 3. Find f(10). Show the key steps, then give the final number."
)
messages = [
{"role": "system", "content": "You are a careful math assistant."},
{"role": "user", "content": problem},
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
enable_thinking=True
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
outputs = model.generate(**inputs, max_new_tokens=4096)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
parsed = processor.parse_response(response, prefix=inputs["input_ids"])
print(parsed)

Set max_new_tokens high when thinking is on. The reasoning comes before the answer and counts toward the limit.

Document and image extraction in a loop

The 31B has the lowest average edit distance on OmniDocBench 1.5 in the table (0.131), which makes it a good fit for document parsing and OCR. The card lists document/PDF parsing, chart comprehension, multilingual OCR and handwriting recognition among its image capabilities. The script below runs the same extraction prompt on each image URL you pass on the command line, one generate call per image. This is a plain loop, not batched inference. The card has no sample document image, so pass URLs to your own scans, receipts or charts. Put the image before the text, as the card recommends.

extract_loop.py
import sys
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "google/gemma-4-31B-it"
# Usage: python extract_loop.py <image_url> [<image_url> ...]
# Pass URLs to your own document, receipt or chart images.
IMAGE_URLS = sys.argv[1:]
if not IMAGE_URLS:
sys.exit("Usage: python extract_loop.py <image_url> [<image_url> ...]")
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)
PROMPT = (
"Extract all readable text from this image. Then return a JSON object with keys "
"'title', 'key_values' (a list of {label, value} pairs) and 'summary' (one sentence). "
"Return only the JSON."
)
results = {}
for url in IMAGE_URLS:
messages = [
{
"role": "user", "content": [
{"type": "image", "url": url},
{"type": "text", "text": PROMPT}
]
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
outputs = model.generate(**inputs, max_new_tokens=1024)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
results[url] = processor.parse_response(response, prefix=inputs["input_ids"])
for url, parsed in results.items():
print(url)
print(parsed)
print("-" * 40)

The model accepts images at any aspect ratio, so you don’t need to crop scans to a square first.

Summarizing short videos

The model reads video as a sequence of frames. The card gives a limit of 60 seconds at one frame per second. That’s enough for short clips like product demos, screen recordings or security footage you want described in text.

video_summary.py
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "google/gemma-4-31B-it"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)
messages = [
{
"role": "user",
"content": [
{"type": "video", "video": "https://github.com/bebechien/gemma/raw/refs/heads/main/videos/ForBiggerBlazes.mp4"},
{"type": "text", "text": "Describe this video as a timeline: list the main scenes in order, then give a one-sentence summary."}
]
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(processor.parse_response(response, prefix=inputs["input_ids"]))

Gotchas

  • No audio on this size. Only E2B, E4B and 12B accept audio input. For speech recognition or speech translation, use one of those models.
  • Memory. This is a 30.7B dense model, so it uses every parameter for every token. The card puts it in the “consumer GPUs and workstations” tier but gives no VRAM figure. device_map="auto" spreads the weights across the devices you have. Check your own hardware before you commit to it.
  • Sampling settings. The card recommends temperature=1.0, top_p=0.95 and top_k=64 for all use cases.
  • Thinking tags still appear when thinking is off. On this model, turning thinking off still produces an empty block, <|channel>thought\n<channel|>, before the answer. Use parse_response instead of slicing strings yourself. If you write your own template, you turn thinking on by putting <|think|> at the start of the system prompt.
  • Multi-turn history. When you add the model’s earlier turns to the conversation, include only the final answers, not the thoughts. The exception is tool-call turns, where you should keep the thinking.
  • Modality order. The card’s best-practices section says to put images before the text (and audio after it). The card’s video example also puts the video before the text, per its code comment, so the scripts here do the same.
  • Visual token budget. Images can use a budget of 70, 140, 280, 560 or 1120 tokens. Use the higher budgets for OCR and small text, and the lower ones for captioning, classification and video. The card doesn’t show which argument sets the budget, so check the processor docs.
  • Video length. The limit is 60 seconds at 1 fps.
  • Knowledge cutoff. The training data ends in January 2025. The card notes the model may generate incorrect or outdated factual statements.
  • License. Apache 2.0, with Google’s Gemma 4 license page linked from the card. Read it before you ship.

When to pick it

Pick the 31B when you want the strongest results in the Gemma 4 family and can afford a dense model of this size. It leads the family on every text, vision and long-context benchmark in the card’s table. The biggest gaps over the 26B A4B are long-context retrieval (66.4% vs 44.1% on MRCR v2 at 128k), Codeforces ELO (2150 vs 1718) and HLE without tools (19.5% vs 8.7%). It’s a strong option for self-hosted document parsing, coding assistants and reasoning-heavy pipelines where the data has to stay on your hardware.

Skip it when:

  • Speed matters more than the last few points. The 26B A4B MoE activates only 3.8B parameters per token. The card says it runs “almost as fast as a 4B-parameter model”, and it stays close on most benchmarks.
  • You need audio. Use the 12B Unified, E4B or E2B.
  • You’re targeting phones or laptops. E2B and E4B are built for on-device use.
  • You need video longer than a minute. You’ll have to split it into clips or sample frames yourself.

The card lists function calling as natively supported but gives no example. Check the Gemma documentation for the tool-calling format before you build an agent on it.

Related