← Writing

codex-img 0.7 and 0.8 add camera angles, exact palettes and a way to review images

Frame shapes that the backend actually follows, camera presets for game art, named styles and characters, palettes snapped to exact colours, then projects with a history of every run, comments that become edits, reruns, and batch versions that are never lost.

On this page

codex-img 0.7 is about the parts of a prompt that kept being written again for every image: the shape, the camera angle, the art style, what a character looks like, and which colours to use. Each one now has an option, and most of them have presets that a project defines once. Claude Code built it, and tested every option on real requests, about 70 images in total, before writing down what worked. Every claim about 0.7 below comes from those tests.

0.8 is about what happens after an image is made: looking at it, saying what's wrong, and getting a better version without losing the one you had. It's further down.

The prompt sets the shape, not --size

Earlier versions passed --size to the backend and documented it as a hint. In tests it wasn't even that: a portrait and a landscape request came back at the same 1312x1199. What does set the shape is a sentence at the start of the prompt, which -a adds:

sh
codex-img "a red apple on a wooden table" -a 16:9 -o apple.png
Four photos of the same red apple on a wooden table, each drawn at its real proportions. The first two, requested with -s 1024x1536 and -s 1536x1024, both came back almost square at 1312x1199. The third, requested with -a 2:3, is a tall 1024x1536, and the fourth, with -a 16:9, a wide 1672x941

The backend keeps about 1.57 megapixels whatever the ratio, so 2:3 lands on exactly 1024x1536. codex-img warns when an image comes back more than 2% off the ratio, and refuses ratios past 3:1, which the backend clamps. It works on edits too: a 2:3 photo edited with -a 16:9 came back 16:9, with the scene widened around the subject. The idea and the pixel counts came from a pull request to another tool.

Camera angles

The racing game's prompts had the camera angle written into every one, in a dozen slightly different ways. Some said "front view at a slight angle", which shows a little of the top surface, and in a game with a low camera that looks like the ground sloping up behind the building. --view replaces all of that with five built-in angles:

A car, a cottage and an oak tree, each drawn in five views. Side shows the car's profile and a flat elevation of the cottage. Front shows the car's nose and the cottage's front wall. Top-down shows the car's roof from directly above and the cottage almost entirely as roof, with a sliver of its front wall. Three-quarter looks down at the front of each, with the roof filling the top half. Isometric turns the car and the cottage 45 degrees, showing the top and two sides
  • side and front are flat elevations with no top surface and a straight bottom edge, for pseudo-3D roadsides and side-scrollers.
  • top-down looks straight down. Cars and trees come out as clean plan views, but buildings keep a sliver of their front wall.
  • three-quarter looks down at about 45° with the front square to the camera, like a classic RPG.
  • isometric turns the subject 45° and shows the top and two sides.

Two of them needed rewording before they worked. The first three-quarter turned the cottage 45°, which is isometric, and the first top-down gave a front view under a big roof:

The same cottage generated with two wordings of each view. Under three-quarter, the first wording turned the cottage 45 degrees like an isometric drawing, and the built-in wording keeps its front facing the camera with the roof above. Under top-down, the first wording showed mostly the front wall under the roof, and the built-in wording shows mostly roof

One more change came from making the picture above. The three-quarter wording said "as in 16-bit RPGs", and twice in a row that turned the car into pixel art: a camera preset was also setting the style. 0.7.1 drops the phrase, and the car in that column was made with the new wording.

Presets a project defines once

Views, styles, characters and palettes are all presets. A project keeps its own in codex-img.json, and codex-img presets lists everything available and where each one is defined:

json
{
  "views": {"roadside": "seen straight on from the side at eye level, its bottom edge a straight line"},
  "styles": {"harbour": {"text": "bright cartoon harbour game art, bold colours", "refs": ["refs/boat.png"]}},
  "characters": {"captain": {"text": "a stocky walrus sea captain with big tusks, in a yellow raincoat", "refs": ["art/captain.png"]}},
  "palettes": {"harbour": "#2B1D14 #6B3E26 #C7743A #F2C14E #F7EBD0 #3B6E5A"}
}
sh
codex-img presets add character captain --ref art/captain.png --text "a stocky walrus sea captain ..."
codex-img presets promote character captain     # copy it to the global presets, for every project
codex-img "the captain waving from a pier" --character captain --view side -a 3:2 -o wave.png

A batch spec can define its own too, and then the spec wins over the project, the project over global presets, and those over the built-ins. Defining a preset with a built-in name replaces the built-in one.

What actually gets sent

Every option that adds to the prompt adds a sentence in a fixed place, and --json shows the whole result as submittedPrompt. In order:

  1. --aspect 3:2: “The frame must be in 3:2 landscape format, wider than it is tall.”
  2. --view side: “Camera: seen perfectly straight on from the side at eye level: a flat side elevation…”
  3. One line per reference image, numbered after any -i images. --character captain adds “Image 1: character reference for "captain": keep the same character (face, proportions, outfit, colours) in a new pose and scene.”, and --composition-ref adds “Image 2: composition reference only: follow its layout and framing, not its subject or style.”
  4. --character captain: “The character "captain": a stocky walrus sea captain with big tusks…”
  5. Your prompt: “The captain waving from the end of a pier.”
  6. --style harbour: the style's text, “Bright cartoon harbour game art, bold colours.”
  7. --palette harbour: “Use only these 6 colours, exactly, and no others: #2B1D14, #6B3E26…”

Reference images with a role

-i sends an image without saying what it's for, so the prompt had to explain it. Three new options say it for you: --style-ref, --composition-ref and --character-ref each add a line telling the model how to use that image.

Three rows, each a reference image, an arrow labelled with the option and the prompt, and the result. A street lamp drawn with flat colours and dark outlines, used with --style-ref and "a red fox sitting", gave a fox in the same flat outlined style. An isometric cottage used with --composition-ref and "a glass greenhouse" gave a greenhouse at the same isometric angle, even with a chimney and front steps. A photo of an apple on a wooden table used with --style-ref gave a photo of a fox on that same table

The last row is the catch. Any input image makes the request an edit, so the result takes the reference's frame and tends to keep its setting: the fox ended up on the apple's table. Describe the new background, and pass -a for another shape. In two A/B tests, the labels made no visible difference compared with a plain -i. They're there so that several images can't be confused.

The same character in every scene

A character preset is an image of the character and a text describing them. The text should describe only their looks. The first version of presets add --from took the anchor image's whole prompt as the text, and that prompt said "isolated on a transparent background". Every scene with the captain then faded to transparent at the edges:

Three images of the same cartoon walrus captain in a yellow raincoat and a white captain's hat. On the left, the anchor image on a transparent background. In the middle, the captain steering a ship's wheel in a storm, with the scene fading out at the edges. On the right, the same storm scene filling the frame

With only his looks as text, the scene fills the frame and the captain keeps his face, coat and hat. --from now takes the image only.

Palettes with exact colours

A game with a fixed palette needs every pixel in it. In tests, asking for that in words didn't work: "a limited palette of 8 colours" changed nothing, and naming a palette without its colours ("PICO-8") put 8% of pixels near it. Listing the hex codes worked much better, with 81–88% of pixels within a small distance of a palette colour. Still, every image had 9,000 to 14,000 distinct colours.

So --palette does both. It adds the hex codes to the prompt, and then it snaps every colour to the nearest palette colour and writes a palette PNG with exactly those colours:

sh
codex-img "a treasure chest sprite" --palette '#2B1D14,#6B3E26,#C7743A,#F2C14E,#F7EBD0,#3B6E5A' -b transparent -o chest.png
codex-img "a treasure chest sprite" --palette pico-8 -b transparent -o chest.png
The same treasure chest prompt with four palettes, each with its colours shown below it: six warm browns and golds with a teal, the four Game Boy greens, the sixteen PICO-8 colours, which turned the chest into pixel art, and the sixty-four colours of Resurrect 64

The model designs around the colours it gets: the Game Boy chest is all greens, and the PICO-8 one became pixel art without being asked. After the snap, these images look almost the same as before it, because the model already kept close to the palette. With 64 colours, the prompt mattered less (33–49% of pixels close), but a 64-colour palette covers most shades, so the snap still looks right.

Art that wasn't made with the palette is harder. Its shading falls between palette colours, and PICO-8 has no dark brown:

Three close-ups of the corner of a wooden treasure chest. As generated, the wood is smooth brown. Snapped with --palette pico-8, the wood turns into purple streaks with dark blue speckles. With --palette-clean added, the same wood is a flat neutral grey with clean edges

--palette-clean first reduces the image to twice the palette's colours, then matches hue before lightness and removes stray pixels. The brown becomes the palette's neutral grey instead of purple. It's off by default, because on art generated with the palette it changes more than it needs to.

There are 14 built-in palettes, taken from Lospec. presets add palette reads your own from hex codes, a GIMP .gpl or .hex file, or a swatch image:

The fourteen built-in palettes as rows of colour swatches: pico-8 with 16 colours, game-boy with 4, nes with 55, c64 with 16, zx-spectrum with 15, cga with 4, ega with 16, ega-64 with 64, sweetie-16 with 16, dawnbringer-16 and dawnbringer-32, endesga-32 with 32, and resurrect-64 and aap-64 with 64 each

A record of how each image was made

batch now writes a manifest beside every raw image, with the prompt that was sent, the presets and where they came from, every input image with a fingerprint, and what the backend reported. When the spec changes later, the asset's line says so:

text
skip   generate crane (raw image exists)
skip   generate captain-wave (raw image exists; note: changed since its raw image was generated, delete it to re-roll)

It never generates the image again by itself, because that spends quota. The same record is written for single images with --manifest. There's no seed, so it can't recreate the exact pixels, but it shows what to change.

Also in 0.7

  • The Agent Skill is half its old size. It keeps what every image request needs, and points to separate references for game art, converting images and the Python fallback, so an agent making an icon doesn't read about sprite pipelines.
  • The prompt guide took in what applies from the image skill that ships with Codex: a labelled-line format for longer prompts, how much detail to add to a vague request, and recipes for slides, wireframes and character consistency.
  • A review of the release found four problems, all fixed before it shipped. A batch could overwrite a reference image with its own output, two presets could share a file that removing one of them deleted, the Python fallback sent presets as a prompt, and changing an asset's background didn't count as a change.

Projects keep a record of every run

A project is a folder with a codex-img.json, the file that already holds its presets, and codex-img init creates an empty one. Every run inside it gets its own file in .codex-img/runs/, one line per event: what was asked for, the backend's progress, where the image was saved, and why a job failed. Paths in it are relative to the project, so the folder can be moved or committed. Every image generated in a project also gets the manifest from 0.7, without --manifest.

Each run also records whether it called the backend, so the images that used quota can be counted from these files, separately from local work like conversion and contact sheets. "events": false in codex-img.json or CODEX_IMG_EVENTS=off turns the run files off, and manifests are still written, because the commands below read them.

Run the same request again

The backend has no seed, so an image can't be made again exactly. The next best thing is sending the same request:

sh
codex-img rerun art/fox.png -n 3 -o art/variations/

rerun reads the image's manifest and sends the prompt that was actually submitted, with its references and settings, then converts the result the way the original was converted. It uses the old preset text even if the preset has changed since. A reference image that has changed since stops it before any quota is spent, and --anyway sends it anyway. The new images record the original as their parent.

Comments that become edits

Reviewing a set of images used to mean keeping notes somewhere else. Now a comment or a star goes into the image's manifest, or onto the asset in a batch spec:

sh
codex-img comments art/shroom.png --set "Give it a small green leaf on top of the cap"
codex-img stars art/hero.png --set true
codex-img comments --json        # every comment in the project

Only that field is changed, and the rest of the file keeps its layout, so in a hand-written spec under git the change is one line. refine turns the comment into an edit of the saved image:

sh
codex-img refine art/shroom.png --from-comment -o art/shroom-leaf.png
Two pixel-art mushroom creatures with a red cap with white spots and big black eyes. The first is the original, with the comment "Give it a small green leaf on top of the cap". The second, made with codex-img refine --from-comment, is the same mushroom with a small green leaf on top of its cap

It sends "Change only: … Keep everything else exactly as it is." with the image, and keeps the image's presets and settings. If a preset has changed since the image was made, it warns and names it, because the edit will use the new text. The comment is cleared only after the new image and its manifest are saved, so a failed edit keeps it.

Batch versions that aren't lost

In 0.7, re-rolling a batch asset meant deleting its raw image. Now there's an option for it, and the old image is kept:

sh
codex-img batch art/assets.json --reroll hero --dry-run   # what it would cost
codex-img batch art/assets.json --reroll hero --expect-images 1
codex-img batch art/assets.json --restore hero .codex-img/history/hero/<version>.png

The current raw image stays in place until the new one is saved, and then moves to .codex-img/history/hero/. If generating fails, nothing moves. --restore puts any version back, or any other image, and costs nothing. --expect-images stops before spending quota if the run would generate a different number of images than the dry run said, because a reference it depends on went missing in between, for example. Two batch runs on the same spec wait for each other, and --no-wait fails instead.

batch --inspect --json lists every asset with its comment, star, whether its spec has changed since it was generated, and its earlier versions. It needs no login.

Checking a spec before it costs anything

A typo in a batch spec used to show up when the batch ran. codex-img check reads a spec or codex-img.json with the same code that runs it, and lists every problem it finds: unknown fields, presets that don't exist, references that point at each other. It needs no login. The repository also has JSON Schemas for both files, and pointing $schema at them gives autocomplete in an editor.

Also in 0.8

  • --nearest resizes by copying pixels, for scaling pixel art up by a whole number, a 32×32 sprite to 96×96 for example.
  • convert --mask-out writes a black and white image of which pixels the background removal took out, and convert --json reports how many pixels hard alpha and background removal changed.
  • Local image processing (conversion, palettes, contact sheets and tiles) moved into a separate library with no network or login code, and convert's output is checked byte for byte against saved files.

Try it

Download the binary for macOS, Linux or Windows from the release page, and log in once with codex login using your ChatGPT account. Run codex-img presets to see the views and palettes, and codex-img init in a game's folder to keep a record of its art. The Agent Skill in the repository tells Claude Code and Codex when to use each of them.