Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Photoproduct - AI Agent, Image Generation, Creative AI

Give Your AI Agent 1,241 Models (In One Integration)

Your AI agent can read files. It can write code. It can call APIs. But it cannot make a picture. Connect it to Photoproduct.io and that changes: one endpoint, seven tools, 1,241 models. Here is a real five step workflow that produced images, video and music end to end for 15 credits.

TL;DR
  • Seven MCP tools and 1,241 models behind one endpoint, priced before you spend
  • Jobs chain by job_ref, so no files ever move through your agent
  • Full five step workflow in this post: images, video and music for 15 credits

Your AI agent can read files. It can write code. It can call APIs.

But it cannot make a picture.

That changes the moment you connect it to Photoproduct.io: one endpoint, seven tools and 1,241 models, all callable the same way your agent already calls everything else.

And in this post I am going to show you exactly what that looks like. Including a real five step workflow that produced the images, the video and the music you will see below, start to finish, for 15 credits ($1.50).

Let's dive in.

What Your Agent Gets Access To

Connect with a bearer token and seven tools appear.

That is the entire surface. No SDK to learn. No per model integration to write.

ToolWhat it does
list_modelsSearch 1,241 models, with prices and reliability data
runStart a generation, get a job_ref back
get_jobPoll it, get the finished file
upload_fileSend a small file inline
upload_from_urlPoint it at a hosted file instead
prompt_guideThe house prompt grammar, per job type
balanceCredits remaining

Here is the part most people miss:

The catalog prices every call before you make it.

Ask list_models for a model and it hands back the credit cost for a typical request. It also takes a max_credits filter.

Which means your agent can shop to a budget. "Show me video models under 10 credits" is a real query that returns a real shortlist.

It gets better. Every entry also carries success_rate and p50_seconds, measured from how that model has actually behaved recently.

Your agent can pick the model that works and finishes on time, instead of guessing from a name.

1,241 Models, One Integration

The catalog is broader than "image generator" suggests:

  • Text to image
  • Image to image
  • Text to video
  • Image to video
  • Video to video
  • Text to speech, and speech to text
  • Text to music

Plus the long tail. meshy/rigging rigs a humanoid 3D model from a GLB URL, which is not something you expect to find next to a text to image endpoint.

Bottom line: one integration, and your agent reaches all of it.

Consistency, Built In

Here is where it gets genuinely interesting.

Making one good image is easy. Making twenty that clearly show the same character is the thing that turns single images into a body of work.

Photoproduct.io ships that as a method, not as luck.

The prompt_guide tool is where it lives. It is not documentation. It is a production pipeline with a defined order:

story_bible -> face_lock -> character_outfit -> character_sheet -> video_from_image -> video_keyframes

Each stage hands back a grammar, a worked example and the models to use for it.

So instead of inventing a process, your agent follows one.

The Five Step Workflow (With Real Output)

Enough theory. Here is the actual chain, run end to end.

Step 1: Lock The Face

One image whose only job is the face.

The grammar is twelve points long, and most of it is about taking things away: a flat neutral gray field, shadowless light, no rim light, no depth of field.

Deliberately plain. That is the point. Everything downstream copies this.

Face lock reference image: a woman with a dark auburn chin-length bob on a flat gray background, lit without shadows, wearing a plain black camisole
Step 1. flux-2-pro. 1 credit.

Step 2: Add The Wardrobe

Now the useful part.

This is an edit, not a new generation. And the input is simply the previous job:

{
  "model": "flux-2-pro/edit",
  "params": {
    "image_urls": ["job_NeOlT6PjpchlYwtG1KvV5w"],
    "image_size": "portrait_4_3",
    "prompt": "Keep this exact person... change only the wardrobe..."
  }
}

Look closely at image_urls. That is a job_ref from step 1, passed straight in.

No downloading. No re-uploading. No files moving through your agent at all.

The same woman, same face and hair, now wearing a charcoal technical half-zip jacket over a black crew neck, on the same flat gray background
Step 2. flux-2-pro/edit. 1 credit.

Step 3: Build The Character Sheet

Three panels in one frame: the garment from the front, the garment from behind, and a tight portrait that re-anchors the face inside the new look.

Three-panel character sheet showing the jacket from the front, the jacket from behind, and a close portrait of the woman wearing it
Step 3. Everything downstream references this instead of re-deriving the look.

Step 4: Make It Move

Kling 2.5 Turbo Pro, image to video, built from the step 2 plate.

8 credits bought 5.04 seconds at 1244x1660 and 24fps.

The grammar here flips the usual advice: describe only what moves, and let the plate carry the appearance. Then name what the brows, the eyelids and the corners of the mouth do on each beat.

That last detail is what makes a take feel performed rather than animated.

Step 4. Generated from the still above.

Step 5: Add Sound

Different category. Same single tool call.

Stable Audio 2.5, 20 seconds of instrumental bed, 4 credits, from a one paragraph brief.

Step 5. stable-audio-25. 4 credits.

The Result

Here is the step 1 face lock next to the final frame of the step 4 video.

Three generations apart.

Side by side comparison: the original face lock on the left and the last frame of the generated video on the right, showing the same face, hair and beauty mark
Same bone structure. Same eyes. Same mark on the same cheek.

Same person.

And that is the whole promise, demonstrated rather than claimed.

What It Cost

The entire workflow above:

StepModelCredits
1. Face lockflux-2-pro1
2. Outfitflux-2-pro/edit1
3. Character sheetflux-2-pro/edit1
4. Videokling-video/v2.5-turbo/pro8
5. Audiostable-audio-254
Total15

15 credits. $1.50. About six minutes.

Credits are $0.10 each, and the quotes you see in list_models are for each model's default request. Pass bigger inputs and the number moves, which is exactly why the budget filter exists.

So query the quote first, then run. One cheap round trip and your agent never surprises you.

Who This Is For

You will get the most out of this if you are building something where media is a step rather than the product:

  • A content workflow that needs images on demand
  • A storyboard or previsualisation tool
  • A game asset pass
  • Any agent pipeline where generations are plural and have to agree with each other

The chaining, the budget filter and the prompt grammar all compound with volume.

Already selling physical products? There is a separate walkthrough aimed at product photography.

How To Try It

Three steps:

  1. Sign in with Google at photoproduct.io and mint a token
  2. Add the MCP server to whichever agent you already use
  3. Ask it one question before you generate anything

Make that question: "Which video models cost 10 credits or less?"

Your agent will call list_models with max_credits and come back with real models and real prices.

That single round trip tells you more than any screenshot will. And it costs nothing.

If you want the wiring details for an MCP server that is not in a registry, this walkthrough covers it.

Workflow run on 2026-10-01 against the photoproduct.io MCP server over streamable HTTP, protocol 2025-06-18. The tool list, model count and prices were read from the live server. Every image, the video and the audio on this page are the actual outputs of the five calls described, unretouched and uncropped except for scaling.

Prev Article
Ruflo
Next Article
Pinokio Computer

Related to this topic: