← Writing

codex-img 0.7 adds camera angles, characters and exact palettes

Frame shapes that the backend actually follows, camera presets for game art, named styles and characters, reference images with roles, palettes snapped to exact colours, and a record of how every image was made.

codex-img 0.7 is about the parts of a prompt that kept being written again for every image: the shape, the camera angle, the art style, what a character looks like, and which colours to use. Each one now has an option, and most of them have presets that a project defines once. Claude Code built it, and tested every option on real requests, about 70 images in total, before writing down what worked. Every claim below comes from those tests.

The prompt sets the shape, not --size

Earlier versions passed --size to the backend and documented it as a hint. In tests it wasn't even that: a portrait and a landscape request came back at the same 1312x1199. What does set the shape is a sentence at the start of the prompt, which -a adds:

sh
codex-img "a red apple on a wooden table" -a 16:9 -o apple.png
Four photos of the same red apple on a wooden table, each drawn at its real proportions. The first two, requested with -s 1024x1536 and -s 1536x1024, both came back almost square at 1312x1199. The third, requested with -a 2:3, is a tall 1024x1536, and the fourth, with -a 16:9, a wide 1672x941

The backend keeps about 1.57 megapixels whatever the ratio, so 2:3 lands on exactly 1024x1536. codex-img warns when an image comes back more than 2% off the ratio, and refuses ratios past 3:1, which the backend clamps. It works on edits too: a 2:3 photo edited with -a 16:9 came back 16:9, with the scene widened around the subject. The idea and the pixel counts came from a pull request to another tool.

Camera angles

The racing game's prompts had the camera angle written into every one, in a dozen slightly different ways. Some said "front view at a slight angle", which shows a little of the top surface, and in a game with a low camera that looks like the ground sloping up behind the building. --view replaces all of that with five built-in angles:

A car, a cottage and an oak tree, each drawn in five views. Side shows the car's profile and a flat elevation of the cottage. Front shows the car's nose and the cottage's front wall. Top-down shows the car's roof from directly above and the cottage almost entirely as roof, with a sliver of its front wall. Three-quarter looks down at the front of each, with the roof filling the top half. Isometric turns the car and the cottage 45 degrees, showing the top and two sides
  • side and front are flat elevations with no top surface and a straight bottom edge, for pseudo-3D roadsides and side-scrollers.
  • top-down looks straight down. Cars and trees come out as clean plan views, but buildings keep a sliver of their front wall.
  • three-quarter looks down at about 45° with the front square to the camera, like a classic RPG.
  • isometric turns the subject 45° and shows the top and two sides.

Two of them needed rewording before they worked. The first three-quarter turned the cottage 45°, which is isometric, and the first top-down gave a front view under a big roof:

The same cottage generated with two wordings of each view. Under three-quarter, the first wording turned the cottage 45 degrees like an isometric drawing, and the built-in wording keeps its front facing the camera with the roof above. Under top-down, the first wording showed mostly the front wall under the roof, and the built-in wording shows mostly roof

One more change came from making the picture above. The three-quarter wording said "as in 16-bit RPGs", and twice in a row that turned the car into pixel art: a camera preset was also setting the style. 0.7.1 drops the phrase, and the car in that column was made with the new wording.

Presets a project defines once

Views, styles, characters and palettes are all presets. A project keeps its own in codex-img.json, and codex-img presets lists everything available and where each one is defined:

json
{
  "views": {"roadside": "seen straight on from the side at eye level, its bottom edge a straight line"},
  "styles": {"harbour": {"text": "bright cartoon harbour game art, bold colours", "refs": ["refs/boat.png"]}},
  "characters": {"captain": {"text": "a stocky walrus sea captain with big tusks, in a yellow raincoat", "refs": ["art/captain.png"]}},
  "palettes": {"harbour": "#2B1D14 #6B3E26 #C7743A #F2C14E #F7EBD0 #3B6E5A"}
}
sh
codex-img presets add character captain --ref art/captain.png --text "a stocky walrus sea captain ..."
codex-img presets promote character captain     # copy it to the global presets, for every project
codex-img "the captain waving from a pier" --character captain --view side -a 3:2 -o wave.png

A batch spec can define its own too, and then the spec wins over the project, the project over global presets, and those over the built-ins. Defining a preset with a built-in name replaces the built-in one.

What actually gets sent

Every option that adds to the prompt adds a sentence in a fixed place, and --json shows the whole result as submittedPrompt. In order:

  1. --aspect 3:2: “The frame must be in 3:2 landscape format, wider than it is tall.”
  2. --view side: “Camera: seen perfectly straight on from the side at eye level: a flat side elevation…”
  3. One line per reference image, numbered after any -i images. --character captain adds “Image 1: character reference for "captain": keep the same character (face, proportions, outfit, colours) in a new pose and scene.”, and --composition-ref adds “Image 2: composition reference only: follow its layout and framing, not its subject or style.”
  4. --character captain: “The character "captain": a stocky walrus sea captain with big tusks…”
  5. Your prompt: “The captain waving from the end of a pier.”
  6. --style harbour: the style's text, “Bright cartoon harbour game art, bold colours.”
  7. --palette harbour: “Use only these 6 colours, exactly, and no others: #2B1D14, #6B3E26…”

Reference images with a role

-i sends an image without saying what it's for, so the prompt had to explain it. Three new options say it for you: --style-ref, --composition-ref and --character-ref each add a line telling the model how to use that image.

Three rows, each a reference image, an arrow labelled with the option and the prompt, and the result. A street lamp drawn with flat colours and dark outlines, used with --style-ref and "a red fox sitting", gave a fox in the same flat outlined style. An isometric cottage used with --composition-ref and "a glass greenhouse" gave a greenhouse at the same isometric angle, even with a chimney and front steps. A photo of an apple on a wooden table used with --style-ref gave a photo of a fox on that same table

The last row is the catch. Any input image makes the request an edit, so the result takes the reference's frame and tends to keep its setting: the fox ended up on the apple's table. Describe the new background, and pass -a for another shape. In two A/B tests, the labels made no visible difference compared with a plain -i. They're there so that several images can't be confused.

The same character in every scene

A character preset is an image of the character and a text describing them. The text should describe only their looks. The first version of presets add --from took the anchor image's whole prompt as the text, and that prompt said "isolated on a transparent background". Every scene with the captain then faded to transparent at the edges:

Three images of the same cartoon walrus captain in a yellow raincoat and a white captain's hat. On the left, the anchor image on a transparent background. In the middle, the captain steering a ship's wheel in a storm, with the scene fading out at the edges. On the right, the same storm scene filling the frame

With only his looks as text, the scene fills the frame and the captain keeps his face, coat and hat. --from now takes the image only.

Palettes with exact colours

A game with a fixed palette needs every pixel in it. In tests, asking for that in words didn't work: "a limited palette of 8 colours" changed nothing, and naming a palette without its colours ("PICO-8") put 8% of pixels near it. Listing the hex codes worked much better, with 81–88% of pixels within a small distance of a palette colour. Still, every image had 9,000 to 14,000 distinct colours.

So --palette does both. It adds the hex codes to the prompt, and then it snaps every colour to the nearest palette colour and writes a palette PNG with exactly those colours:

sh
codex-img "a treasure chest sprite" --palette '#2B1D14,#6B3E26,#C7743A,#F2C14E,#F7EBD0,#3B6E5A' -b transparent -o chest.png
codex-img "a treasure chest sprite" --palette pico-8 -b transparent -o chest.png
The same treasure chest prompt with four palettes, each with its colours shown below it: six warm browns and golds with a teal, the four Game Boy greens, the sixteen PICO-8 colours, which turned the chest into pixel art, and the sixty-four colours of Resurrect 64

The model designs around the colours it gets: the Game Boy chest is all greens, and the PICO-8 one became pixel art without being asked. After the snap, these images look almost the same as before it, because the model already kept close to the palette. With 64 colours, the prompt mattered less (33–49% of pixels close), but a 64-colour palette covers most shades, so the snap still looks right.

Art that wasn't made with the palette is harder. Its shading falls between palette colours, and PICO-8 has no dark brown:

Three close-ups of the corner of a wooden treasure chest. As generated, the wood is smooth brown. Snapped with --palette pico-8, the wood turns into purple streaks with dark blue speckles. With --palette-clean added, the same wood is a flat neutral grey with clean edges

--palette-clean first reduces the image to twice the palette's colours, then matches hue before lightness and removes stray pixels. The brown becomes the palette's neutral grey instead of purple. It's off by default, because on art generated with the palette it changes more than it needs to.

There are 14 built-in palettes, taken from Lospec. presets add palette reads your own from hex codes, a GIMP .gpl or .hex file, or a swatch image:

The fourteen built-in palettes as rows of colour swatches: pico-8 with 16 colours, game-boy with 4, nes with 55, c64 with 16, zx-spectrum with 15, cga with 4, ega with 16, ega-64 with 64, sweetie-16 with 16, dawnbringer-16 and dawnbringer-32, endesga-32 with 32, and resurrect-64 and aap-64 with 64 each

A record of how each image was made

batch now writes a manifest beside every raw image, with the prompt that was sent, the presets and where they came from, every input image with a fingerprint, and what the backend reported. When the spec changes later, the asset's line says so:

text
skip   generate crane (raw image exists)
skip   generate captain-wave (raw image exists; note: changed since its raw image was generated, delete it to re-roll)

It never generates the image again by itself, because that spends quota. The same record is written for single images with --manifest. There's no seed, so it can't recreate the exact pixels, but it shows what to change.

Also in 0.7

  • The Agent Skill is half its old size. It keeps what every image request needs, and points to separate references for game art, converting images and the Python fallback, so an agent making an icon doesn't read about sprite pipelines.
  • The prompt guide took in what applies from the image skill that ships with Codex: a labelled-line format for longer prompts, how much detail to add to a vague request, and recipes for slides, wireframes and character consistency.
  • A review of the release found four problems, all fixed before it shipped. A batch could overwrite a reference image with its own output, two presets could share a file that removing one of them deleted, the Python fallback sent presets as a prompt, and changing an asset's background didn't count as a change.

Try it

Download the binary for macOS, Linux or Windows from the release page, and log in once with codex login using your ChatGPT account. Run codex-img presets to see the views and palettes, and the Agent Skill in the repository tells Claude Code and Codex when to use them.