Skip to content

feat: add Generate resource for AI image generation (Media Generation API) - #15

Open
adidagancld wants to merge 3 commits into
masterfrom
feat/media-generation
Open

adidagancld wants to merge 3 commits into
masterfrom
feat/media-generation

Conversation

@adidagancld

Copy link
Copy Markdown

Summary

Adds a Generate resource that creates images with AI through the Cloudinary Media Generation API (Image Generation add-on). It has three operations, one per operationId in the service's OpenAPI spec (CloudinaryLtd/media_generation → app/schema/schema.yml, v1.0.0):

Operation Endpoint
Generate Image From Text POST /v2/generate/{cloud}/text_to_image
Generate Image From Reference Images (1–4 URLs or asset IDs) POST /v2/generate/{cloud}/image_to_image
Get Generation Task GET /v2/generate/{cloud}/tasks/{task_id}

Everything is additive: a new resource value, with new fields shown only for it. There's no typeVersion bump, and backward-compatibility-check reports 0 breaking changes. Version goes from 0.2.3 to 0.3.0.

Design

  • Spec-driven client, no runtime deps. mediaGeneration.client.ts is hand-written against the spec. Its types mirror components.schemas one-to-one: each oneOf becomes a TS union, and names stay snake_case as on the wire. The UI option lists (model IDs, families, aspect ratios, …) are built from its constants, so a spec change is edited in one place. I didn't use a code generator because of the no-runtime-deps rule, and the API is only three endpoints.
  • Auth is HTTP Basic, the same as the Admin API flow, but on the /v2 base with JSON bodies.
  • Request mapping. The spec's nested model, image_size and target objects are flattened in the UI into a mode selector plus fields that only show for that mode. Pure builders in operations/generate/shared.ts map them back and validate against the spec's bounds before any request: 64–4096 px, 1–4 references, HTTPS URLs, the task-ID pattern. Only fields the user set are sent, so the service defaults apply to everything else.
  • Model IDs are split per operation. -edit models are only valid on image_to_image and the others only on text_to_image, so each operation has its own Model ID list.
  • Output. Each generated image becomes one item, with storage flattened to the top level. That puts public_id, asset_id, secure_url, resource_type, type and version alongside format, width, height, model, seed, request_id and limits. It pipes straight into Transform ({{ $json.public_id }}) and Asset ({{ $json.asset_id }}). Turning Simplify off returns the raw body.
  • Sync vs. async. The default call waits for the image server-side (200). With Options → Async on, it returns a task_id (202) to poll with Get Generation Task or to receive via a Notification URL webhook. The node never polls on its own.
  • Errors show the spec's error.code, with a hint per status and the request_id. For example: MG_00704: reference asset '…' not found, or on 429, how many quota units remain.

Non-obvious: rewriting the NodeApiError in place

httpRequestWithAuthentication already throws a NodeApiError, and new NodeApiError(node, err) returns err untouched when it already is one, silently ignoring the message/description options. So the client reads the response body from error.context.data, where NodeApiError keeps it, and rewrites message / description / httpCode on the existing error. Without this, users would see n8n's generic "The service is receiving too many requests from you" instead of the API's message. This is covered by a unit test and was verified inside real n8n (below).

Testing

  • Unit. 21 new Vitest tests covering the request contract for each operation, every ModelSelection / ImageSize / Target form, validation, output shaping, and error mapping with a real NodeApiError. The suite is now 263 tests, all passing; lint (both configs) and tsc are clean.

  • Live API smoke (cloud adidagan). I ran all three operations against production: the default model, family/tier, auto model, exact dimensions, temporary vs. managed storage, image-to-image with a URL and an asset-ID reference, async → poll to completed, and the 401 / 400 / 404 errors.

  • Real n8n 2.41.3. I installed the package exactly the way n8n installs a community package and ran fixtures/generate-smoke.workflow.json with n8n execute. All 9 nodes succeeded:

    • expressions inside the Options collection and inside a Reference Images row;
    • Generate output feeding a Transform op;
    • async → Wait node → Get Generation Task;
    • MG_… errors surfaced through n8n's real HTTP helper with continue on fail.

    n8n also generated the AI-tool variant (cloudinaryTool) with the Generate resource.

Local n8n e2e (docs)

I rewrote local-n8n-test.md, which had hardcoded paths from one machine and covered only the 0.0.9 upgrade. It now comes with portable scripts under local-n8n/:

  • an isolated instance that never touches ~/.n8n;
  • install-branch.sh, which mirrors n8n's downloadPackage step. That means no npm link, which would load the repo's own n8n-workflow and break instanceof checks;
  • a credential import that reads the gitignored .local/smoke.env and never prints the secret;
  • a helper that summarizes the last execution;
  • the Generate fixture.

The guide also documents two setup problems. n8n needs Node 24 or newer, and its native build needs a Python that still has distutils (PYTHON=/usr/bin/python3). And n8n execute fails silently while the editor is running, because of the task-broker port.

Notes for reviewers

  • Quota is weighted by model. One default-model (nano-banana-2, premium) generation used 9 units, while FLUX.2 Klein used 1. The 429 hint says "quota units" rather than "generations" for that reason, and the README mentions it.
  • Not included: OAuth2 (the spec allows it, but the node only has the API-key credential) and Image-to-Video, which isn't in this spec.
  • The editor-UI checklist in section 3 of the guide (field visibility, the 4-row Reference Images limit, AI Agent attachment) is worth a look during review.

🤖 Generated with Claude Code

… API)

Adds a Generate resource backed by the Cloudinary Media Generation API
(Image Generation add-on): Generate Image From Text, Generate Image From
Reference Images, and Get Generation Task.

- mediaGeneration.client.ts: dependency-free client whose types mirror the
  OpenAPI spec in CloudinaryLtd/media_generation (app/schema/schema.yml),
  one function per operationId, Basic auth on the /v2 base.
- Errors surface the spec's error code, a per-status hint and request_id.
  httpRequestWithAuthentication already throws a NodeApiError, which
  `new NodeApiError()` returns untouched, so the client rewrites it in place.
- Output is one item per generated image with storage fields flattened
  (public_id, asset_id, secure_url) so it pipes into Transform/Asset ops;
  Simplify off returns the raw response. Async returns task_id for polling.
- Additive only (new resource value + gated fields): no typeVersion bump.
- Local n8n e2e: portable scripts + Generate fixture under
  docs/backward-compatibility-check/local-n8n/, guide rewritten.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sync the client with the latest media_generation schema:
- Add the experimental root-level `notices` array ({severity, text}) to the
  200, 202, task GET and error/429 envelopes.
- Carry notices onto simplified output items alongside limits/request_id.
- Show notices in the error description in place of the generic per-status
  hint (the quota-wall notice says not to retry; the 429 hint said to).
- Model family/tier now report `none` instead of `unmapped` (doc only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@eitanp461

Copy link
Copy Markdown
Contributor

Review — pass 1

Head 36338ba · base 87ba5c0 · 2026-10-06

Verdict: approve with nits. Nothing blocking.

Hunted: async/sync/failed task classification and output; error extraction vs. what n8n actually throws, continueOnFail; auth and task-ID/URL injection; request bodies vs. the OpenAPI spec; backward compat, runtime deps, scanner, packaging; whether tests reach real code paths.

Findings

  1. Failed task returns as a success item — nodes/Cloudinary/operations/generate/shared.ts:252-255. Get Task with status: 'failed' outputs {task_id, status: 'failed', request_id}. Nothing throws, so continueOnFail / the error output never fire, and downstream {{$json.public_id}} is just undefined. Suggest throwing a NodeOperationError on failed, or documenting that users must branch on status.
  2. Item index lost on errors — nodes/Cloudinary/mediaGeneration.client.ts:456-460. The rewritten error doesn't set context.itemIndex, so n8n can't tell which item failed. (Existing ops share this gap.) Suggest error.context.itemIndex = i.
  3. Fallbacks for error shapes n8n doesn't throw — nodes/Cloudinary/mediaGeneration.client.ts:387-396. NodeApiError keeps an object body on context.data; the response.body / response.data / JSON.parse branches cover shapes httpRequestWithAuthentication doesn't produce, and the test "handles a plain error carrying the body under response.body" pins one of them. Suggest reading only context.data and dropping that test.
  4. README overstates 429 — README.md:73. A 429 can also come from rate limiting with a plain-text body; that path fails the apiError?.message check and shows n8n's generic error with no quota numbers. Suggest softening the wording.
  5. (speculative) 202 without data loses the task ID — nodes/Cloudinary/mediaGeneration.client.ts:310. TaskResponse.data is optional in the spec; a 202 without it would be read as a sync result. No known server path emits this.

PR body

  • "21 new Vitest tests" — generate.test.ts has 27.
  • Additive / no typeVersion bump, pre-request validation, only-set-fields-sent, model ID lists matching the spec: confirmed.

Not run

Tests, tsc, and the community scanner locally (CI test passed).

@eitanp461

Copy link
Copy Markdown
Contributor

Head

36338ba0fcc2ccc5e32224c6162284ce96164838 · base 87ba5c07d9216e64addbe1b85577963df456b4c8

Open findings

  1. shared.ts:252-255 — failed task returns as success item · CONFIRMED · NON-BLOCKING · verified-at 2026-10-06 @ 36338ba · raised: pass 1
  2. mediaGeneration.client.ts:456-460 — context.itemIndex not set · CONFIRMED · NON-BLOCKING · verified-at 2026-10-06 @ 36338ba · raised: pass 1
  3. mediaGeneration.client.ts:387-396 — speculative error-shape fallbacks + test pinning one · CONFIRMED · NON-BLOCKING · verified-at 2026-10-06 @ 36338ba · raised: pass 1
  4. README.md:73 — 429 wording overstates quota info · CONFIRMED · NON-BLOCKING · verified-at 2026-10-06 @ 36338ba · raised: pass 1
  5. mediaGeneration.client.ts:310 — 202 without data read as sync · SPECULATIVE · NON-BLOCKING · verified-at 2026-10-06 @ 36338ba · raised: pass 1

Closed

none

Reviewed under

  • pass 1 @ 36338ba: task classification, error extraction/continueOnFail, auth + injection, bodies vs. spec, backward compat/deps/scanner/packaging, test reach

Responses

none yet

Next

Verify each open finding against the new head; pick an unmined angle (e.g. UI field gating/displayOptions, AI-tool usage).

@eitanp461

Copy link
Copy Markdown
Contributor

Credit cost hints: link to live data, don't hardcode it

The Image Generation docs don't publish per-model credit costs. They only say cost varies by model and resolution. Hardcoded numbers will drift from the backend without anyone noticing:

  • README.md:73: "9 for Nano Banana 2 versus 1 for FLUX.2 Klein"
  • docs/backward-compatibility-check/local-n8n-test.md:72: "About 6 add-on quota units per run"

Suggestion:

  1. Drop the per-model numbers and say cost varies by model and resolution.
  2. Add a short hint on the Model field (generate.fields.ts:174) pointing to data that stays current:

    Credit use varies by model and resolution. Each result reports credits used and remaining under limits. Manage your plan on the Image Generation add-on page.

@eitanp461

Copy link
Copy Markdown
Contributor

Model Tier: difference between Standard and Premium is unclear

nodes/Cloudinary/descriptions/generate.fields.ts:208 — Model Tier offers Standard / Premium with no description, so users can't tell what they trade off by picking one over the other. The Image Generation docs don't define it either; they only list which model currently sits in each slot.

@eitanp461

Copy link
Copy Markdown
Contributor

Reuse buildComponents instead of re-implementing it

The "run builder → rethrow as NodeOperationError with itemIndex" wrapper now exists three times:

  • nodes/Cloudinary/operations/transform/shared.ts:234 — buildComponents (existing)
  • nodes/Cloudinary/operations/generate/shared.ts:142 — build, same logic, generic return type
  • nodes/Cloudinary/operations/generate/imageToImage.ts:17-22 — same try/catch inline around referenceImages(rows)

Suggest making buildComponents generic (<T>(ctx, i, build: () => T, prefix = ''): T) and using it in both Generate spots, either imported from transform/shared.ts or moved to cloudinary.utils.ts. Existing Transform callers stay unchanged.

// ── Model selection (`ModelSelection` = ModelByFamily | ModelById | ModelAuto) ──

/** `ModelByFamily.family` */
export const MODEL_FAMILIES = ['flux', 'recraft', 'gpt-image', 'nano-banana', 'ideogram'] as const;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@adidagancld note that this list needs to be maintained and matched with changes that are happening in production.

- Get Generation Task throws on a failed task so Continue On Fail / the
  error output apply
- Set context.itemIndex on every Media Generation API error
- Read the error body only from NodeApiError context.data; drop the
  speculative response.body / JSON.parse fallbacks and their test
- Soften the README 429 wording; drop hardcoded per-model credit costs
- Add a credit-use hint on the Model field and describe Standard vs
  Premium tiers
- Make buildComponents generic and reuse it in the Generate builders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adidagancld

Copy link
Copy Markdown
Author

Thanks @eitanp461 — addressed in 8bf6ec2.

Review pass 1

  1. Failed task as success — fixed. Get Generation Task now throws a NodeOperationError on status: failed (with notices, request_id, and itemIndex), so Continue On Fail / the error output apply. Documented in the README.
  2. itemIndex lost — fixed. context.itemIndex = i is set on every Media Generation error, including the non-spec-body path that keeps n8n's own error.
  3. Speculative error-shape fallbacks — removed. extractMediaGenerationError reads only httpCode and context.data; the response.body test is dropped, replaced by tests for itemIndex and for a plain-text 429 keeping n8n's own error.
  4. README 429 wording — softened: a 429 can also be rate limiting with a plain-text body, in which case n8n's generic error shows.
  5. 202 without data — left as is. No server path emits it, and telling 200 from 202 would mean switching to returnFullResponse, which changes the error and response shapes. Happy to revisit if the spec tightens.

Credit cost hints — dropped the per-model numbers from the README and local-n8n-test.md; added your suggested hint (credits under limits, link to the add-on page) on the Model field.

Model Tier — each option now has a description, taken from the OpenAPI spec since the public docs don't define it: Standard = the family's faster, lower-cost model; Premium = its highest-quality model, typically more credits. The field notes tiers stay valid as models are upgraded.

buildComponents reuse — made generic (<T = string[]>) in transform/shared.ts; both Generate spots use it. Transform callers are unchanged.

Tests (271), build, and lint pass.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants