Module 1 · Foundations of Agentic AI · scripted
Exercise: The Agent's Eyes
The goal: let the agent see what you see
Your first agents drove the conversation and you reported back in words. Now upgrade the channel: in this exercise you answer the agent with photographs. The Agent Lab's input bar has a 📷 button — attach an image from your computer or phone (on a phone, it opens the camera), and it is sent into the conversation as part of your message. The agent sees it.
The exercise at a glance
Run everything in the Agent Lab or your own LLM — attach photos with the 📷 button (on a phone it opens the camera).
Open the Agent Lab → — everything you
built earlier is still there (your saved agents and conversations
included). The Prompts view works too: an attached image shows up
in the log as [image attached], one more thing on the growing
prompt.
Step 1 · Use a picture to get a task done
Step 1 · What you'll do
The picture is the context; the task tells the agent what the context is for. The test of a good pairing: the agent's answer uses things from the image you would never have thought to mention.
Capture — end of Step 1
Step 2 · Use a picture to get feedback on your work
Step 2 · What you'll do
This is the whiteboard move from the lesson, aimed at your own work. The same drawing, prompted three ways, gives you a skeptic, a gap-finder, and a facilitator — and a critique session your group can actually argue with.
Capture — end of Step 2
Step 3 · Turn a picture into a simulation
Step 3 · What you'll do
A drawing of a system is enough for the agent to become the system: a persona-based simulation, generated from a photograph. This is prototyping with a pen — you can test a process or an interface before anything exists.
Capture — end of Step 3
Deliverable
Three captures: the photo-driven task, your group's strongest critique, and a moment from your simulation. We will compare the best of each — and the best paper-drawn system of the session.
- The photos went into the conversation like any other message. What does that mean for the cost of a conversation full of photos — and for what should happen to old ones?
- Where in your research would an agent's eyes replace a measurement you currently type in by hand?