Module 1 · Foundations of Agentic AI · scripted

Exercise: The Agent's Eyes

25 min

The goal: let the agent see what you see

Your first agents drove the conversation and you reported back in words. Now upgrade the channel: in this exercise you answer the agent with photographs. The Agent Lab's input bar has a 📷 button — attach an image from your computer or phone (on a phone, it opens the camera), and it is sent into the conversation as part of your message. The agent sees it.

The exercise at a glance

Use a picture to get a task done.
Use a picture to get feedback on your work.
Turn a picture into a simulation.

Run everything in the Agent Lab or your own LLM — attach photos with the 📷 button (on a phone it opens the camera).

Open the Agent Lab → — everything you built earlier is still there (your saved agents and conversations included). The Prompts view works too: an attached image shows up in the log as [image attached], one more thing on the growing prompt.

Step 1 · Use a picture to get a task done

Step 1 · What you'll do

Take a picture of something around you — your desk, a bookshelf, the contents of a bag, a schedule on your screen.
Send it with a task: "plan the cleanup" · "pick my next three reads and an order" · "what am I forgetting for this trip?" · "turn this schedule into a plan for tomorrow".
Watch for the details it uses that you never typed — that's the picture doing work words would have dropped.

The picture is the context; the task tells the agent what the context is for. The test of a good pairing: the agent's answer uses things from the image you would never have thought to mention.

Capture — end of Step 1

The photo and the task you gave with it.
One thing the agent used from the image that you had not mentioned — and would not have typed.

Step 2 · Use a picture to get feedback on your work

Step 2 · What you'll do

As a group, draw a plan or an idea — on paper or a whiteboard — or screenshot something you're working on.
Photograph it and send it three times, changing the role:
"Act as a skeptic. Poke holes in this — how does it fail in ways we haven't thought of?"
"What are the gaps and ambiguities in this?"
"What are the critical questions we should be asking?"
Answer its follow-up questions and see how the critique sharpens.

This is the whiteboard move from the lesson, aimed at your own work. The same drawing, prompted three ways, gives you a skeptic, a gap-finder, and a facilitator — and a critique session your group can actually argue with.

Capture — end of Step 2

The drawing you photographed.
The strongest single criticism or question the agent produced — the one your group had not thought of.

Step 3 · Turn a picture into a simulation

Step 3 · What you'll do

Draw a process diagram or a user interface sketch and photograph it. (A process from your research works well.)
Prompt: "Act as the system in this image. Tell me how I'm allowed to interact with you, then let me interact with you. Simulate what the system does, with outputs."
Interact with it: step through the process, or "click" around the interface, and see whether the simulation stays true to your drawing.

A drawing of a system is enough for the agent to become the system: a persona-based simulation, generated from a photograph. This is prototyping with a pen — you can test a process or an interface before anything exists.

Capture — end of Step 3

The diagram and your opening prompt.
One exchange from the simulation — and one place where it followed your drawing exactly, or departed from it.

Deliverable

Three captures: the photo-driven task, your group's strongest critique, and a moment from your simulation. We will compare the best of each — and the best paper-drawn system of the session.

As you compare
  • The photos went into the conversation like any other message. What does that mean for the cost of a conversation full of photos — and for what should happen to old ones?
  • Where in your research would an agent's eyes replace a measurement you currently type in by hand?