What is this article talking about?
I often hand the YouTube videos I watch to AI, which automatically organizes them into an opinion report webpage that I keep on my own website to look up anytime. This article is the full demonstration from the free lecture on 2026-06-14: from the video link to the verbatim draft, local organizing, and web page layout, I show you the whole thing, focusing on the thinking behind each step.
· People who often watch instructional videos, talks, and press conferences, forget them right after, and want to keep assets they can trace back to
· People who already know how to use NotebookLM or a desktop Agent and want to connect the two into one process
· People who want to learn how to hand tasks to AI, how to check and accept the results, and how to save a successful workflow
· A complete workflow from video link to opinion report webpage
· Three practical habits for handing tasks to AI: test the limits of its ability, ask it to paraphrase, and give good examples
· Two practical tips for saving Tokens, and a way to distill the process into a skill package
The entire workflow looks like this
When you see a video worth organizing, hand the link to the desktop Agent.
Push the video to NotebookLM to get the text; it's a very powerful YouTube transcript organizer.
Pull the verbatim file back into your own folder and write clear keywords in the file name, such as "subtitle organization."
Organize the key points according to the "opinion report" format you defined in advance.
Use the open-source design skill package to build a webpage, with links that jump back to the matching part of the video.
The whole process is written into a skill package, so next time you can rerun it in one sentence.
Each step looks simple on its own. The value comes from linking them into one line: a video goes from being "watched" to a knowledge asset that is "organized, can be questioned, reviewed, and shared." If you want to see what the finished product looks like first, this opinion report on Musk and knowledge workers is a real example made with this workflow and posted on the official website.
Step one: the source, get the verbatim draft first
For any data you want AI to read, text is still the best form. The easiest way to get a transcript of a video is to drop it into NotebookLM. If adding videos manually each time is a hassle, the browser has a ready-made plug-in that adds a video to NotebookLM with one click on the YouTube page, and it can batch-process an entire channel too.
One step further is to let the desktop Agent automatically handle the whole stretch of "push to NotebookLM, convert the verbatim draft, and pull it back to the local machine." Someone in the community has open-sourced a way to connect an Agent to NotebookLM, but this kind of linkage is an unofficial operation. When I run into a tool like this, I don't install it right away; I first ask the AI three things:
The conclusion after checking: it works, but there's a chance the platform bans it. So my choice was to shrink its use to the minimum and cut every other fancy feature, keeping only "convert the verbatim draft and pull it back" to lower the chance of something going wrong. Remember this is still an unofficial method, and the risk of your account being restricted by the platform can't be reduced to zero. If that bothers you, take the official route of the browser plug-in or manually dropping videos into NotebookLM: slower, but steady. I also agreed with the Agent that this step only uses local commands and tools and doesn't operate the browser screen, because operating the screen directly burns a lot of Tokens.
Three practical habits for handing in tasks
1. Test the limits of its ability first. Don't ask an intern to do a director's job.
The truly hard part of using AI is judging where the ceiling of its ability sits. My approach is to give its ability a rough score: I know this AI has a 100-point ability, so I hand over tasks below 80 points without worry; tasks around 100 points I test first; tasks above 120 points I don't hand over at all, because it can't do them. People ask me how I can comfortably hand work to AI, whether I worry about hallucinations. I only ask it to do what I know it can do, and when I give it clear information, hallucinations naturally drop.
2. Ask it to explain in its own words.
This is just like a boss assigning work to employees. Once I hand over the task, I confirm again: Do you really understand? Explain it back to me in your own words. It works the other way too. When the AI gives me three options, A, B, and C, and I think C looks good, I'll say, "Does C mean this, and what result will it reach?" Once it confirms I've understood correctly, then I pick C.
3. Give a good example when you check the work. If it's ugly, just say it's ugly.
The line "help me format this" actually says nothing. If you're not happy with the result, just say it straight: I find this hard to read, the layout is messy, this key point isn't what I wanted. The most effective move is to give a good example: "I think this version is great, it has a key-point summary up front, the content of each section in the middle, and clicking jumps to the matching part of the video; go in this direction." For the layout foundation you can lean on an open-source design skill package: first have the AI search online for what's free and usable, then fit it into your process.
The opinion report must be defined by you, and the file name must speak for itself
"Let me understand it quickly" means something different to everyone. So I carefully defined things with the AI: what counts as a key point, what a summary is, what an opinion report is, and what format each maps to. A video summary answers "what was said"; an opinion report also answers "why it matters" and "how it can be used." So my opinion report has three layers: source facts, opinion summary, and action suggestions. Once you've made a version you like, tell it: "Remember, from now on this thing is called an opinion report, and whenever I say opinion report, make it like this."
Another little habit that saves a lot of Tokens is putting keywords in the file name. Once you've accumulated a lot of material, there are hundreds of files in a folder, and you say, "Pull up the subtitle organization from the last Apple event." If the file name has no keywords, the AI has to open and search them one by one, burning a pile of tokens just to find the file. Put "subtitle organization" plus a topic in the file name and it searches the file names and hits it directly. I agreed with the Agent: when looking for a file, search the file names first, and only search inside files if that fails.
Distillation: "Remember" must be written into the skill package
When the AI says "I'll remember this," don't rush to believe it; make sure where it's remembered. Stored in the tool's own memory, it disappears the moment you switch to another company; written into the skill package, every Agent can call on it. Use this one if it works well today, switch to another if it's stronger tomorrow, and the script follows you.
And even after it's written, I still press to the end:
Once the process runs smoothly, there's no need to accept someone else's version wholesale. I just looked at other people's open-source linkage methods and asked the AI to cut every feature I didn't use, turning it into a minimalist version that purely grabs verbatim drafts. No gold or silver nest beats your own little doghouse: the workflow and skill package grown out of your own process fit your own needs best. Today's Agents have built-in abilities to create and modify skill packages, so just ask it to change them.
How you can practice: Split the video task into four cards
To turn a video into a report, you can split it into four cards. Write each card's input and output clearly so the Agent can pick up and continue next time.
- Source acquisition: the video link comes in, the verbatim text goes out. Note which parts have public subtitles and which need separate handling.
- Content understanding: the verbatim draft comes in, the key points and context go out. Use NotebookLM or conversational AI to probe the core viewpoints, points of conflict, and target audience.
- Report output: the key points come in, the opinion report webpage goes out. Following the three-layer structure you defined: source facts, opinion summary, and action suggestions.
- Process distillation: this round's method comes in, the skill package and work log go out. Next time a new video shows up, rerun it in one sentence.
Start with a relatively short video and run each of the four cards once. After running through it, you'll find you've learned more than "turning YouTube into a report." You start to know how to design your own workflow and how to train your own AI employees.