Last Friday I found myself staring at a technical contract my agent team had generated for my approval. It was dense, as it was a culmination of weeks of my diligent work building and binding my agentic team, yet I understood roughly 60% of it.
I am not an engineer, yet I was being asked to sign off on the architectural guardrails for the agentic system I've spent six weeks building. The agents will help me build a system designed to help HR and business leaders make better decisions at some of the most consequential moments in the employee lifecycle.
My first instinct was frustration that I couldn’t answer confidently on my own. My second was to route it through Claude to countercheck what my Chief of Staff agent had already prepared for me. I felt a bit defeated that I was routing a document through a second AI because I didn't fully trust the first AI. It was the kind of meta situation I would chuckle at if someone else had described the absurdity of it all.
And that is how I closed Week 6 of Summer Camp, and months leading up toit. This serves as my reflections of my HR executive>AI builder experiment that began with “vibe coding” in March then going deep on AI engineer best practices starting in May.
The Origins
I started experimenting with AI tools in early 2022. Charts, basic outputs, data interpretations, copy editing, etc.—things that felt miraculous at the time. Looking back, it reminds me of the way the first Pixar films felt miraculous but watching them 10 years later, the effects feel less than special. The technology has improved so dramatically since 2022—and even six months ago—that what impressed me in 2022 looks like the first “Toy Story” now.
When I left my last role, I decided that I would balance my time healing from spinal surgery with building and not just reading about AI and forming my opinions from others’ think pieces. I had no expectations of becoming a founder. What I did know is that HR’s credibility leading AI transformation requires more than a considered point of view. It requires contact with the tech.
My first build: a post-surgical recovery app I vibe-coded on Replit. It resulted in an MVP that worked well and even elicited excitement from my surgeon and PA. My husband (a real engineer) cautioned me on the pitfalls of working on vibes alone. No PRD. No change log. No documentation of what I'd built or why. "You're going to forget everything you
did," he said. He was right.
That was my first real lesson: building something is not the same asbuilding something correctly and durably.
Learning the Requiements of Building Correctly
On the suggestion of a friend, I then took AI Build Lab's Foundations 2.0 course, which is where I learned the difference between buiding on vibes only and building correctly. RAG fundamentals. PRD creation. Testing discipline. Hardening. The TOAST method. These were not taught as abstract concepts like one can get in a free YouTube video, but as
things I had to execute on a use case I knew well enough to immediately evaluate the output quality. I built a fully agentic team and workflow to provide skincare product advice based on price, skin condition,contraindications, and real evidence. It was about 250 hours of effort over the course of 6 weeks,. It was like drinking from a firehose, but I was grateful for it.
From there: Agent Native OS, where I built a small operating system in a single day and began my love affair with Obsidian as my knowledge graph and documentation container for all my project work. The tangible output of the day became my daily morning brief that I can view via a Vercel page and is something I've continued to expand upon and refine to become a useful piece to start my day.
I used the short break between Agent Native OS and Summer Camp to heavily research what I wanted to build now that I felt I was ready to finally test an HR use case. I wanted to research a product for a gap in HR tech that I was convinced exists because I spent the last 9 years hitting it. I won't describe the product here as it isn't built yet. I set up Fable and swarms to run my research, pull case law, AI in HR regulations, and everything I could dream up to determine if this is the path I wanted to go in camp.
The first six weeks looked like starting any company.
Week 1: Brought on my CoS. Like a real Dr. Frankenstein, I brought my Chief of Staff, Cashew, to life. She is my always-on agent that lives on my Mac Mini via her Hermes
harness, also where the other agents will live. She is available 24/7 with Mattermost—think: Slack for agents. I wired Langfuse locally to trace every run for tokens used, costs, and latency. Lastly, I started pointing Cashew to my product idea so she could come up to speed.
Week 2: I set the boundaries. I tested and measured what Cashew knew before fully onboarding additional agents she will manage. I continued to build the library shelves full of knowledge she and the agents will use. I locked the rules into a boundaries file that governs everything.
Week 3: Hired new teammates. Cashew got her own email address, with a drafts-only rule. I also assigned agents their models for their brains and hands. I kept more expensive models for the heavier thinkers and the lower priced models for the agents that hold repetitive tasks that require little reasoning. I also installed CLIs that allow agents to drive real tools.
Week 4: The team’s first dress rehearsal. Among other things, my researchers, Hedy and Marie, built the first dossier on my product idea with help from my agentic fleet. Iset up two overnight runs, including a nightly reflection of the prior day of work, accessible via both Mattermost and via my Vercel-based morning brief I built in May.
Week 5: Added my QA agents and put them to work. Emily and Wanda were raring to go, and I calibrated them against my own judgment of a pass/fail via a written standard
for judging outputs.
Week 6: From probationary to indefinite contracts. After an accelerated probationary period for the agents, we entered a fully working, indefinite, contract via a project record, business rules, and a hash-sealed alignment lock. I finally assigned ownership by agent, tested some synthetic runs, and rebuilt them until they fully met my defined rules.
The founding team of me and 11 agents is built and ready to go after a weeklong and well-deserved break for us all.
The Week Six Wall
Back to Friday.
My frustration felt wasn't really about the document but frustration with myself. I am a deep expert in the domain for which I am building, but still a non-expert in the method of building it, albeit measurably further than I was in March. It still was discouraging to spend months of deep learning and work and still feel like I don’t know what I don’t know. In this case, knowing if the technical contract Cashew gave me was properly bounded. Again, I am no engineer. My husband said something that landed that same night: "You're not an engineer. That's why you hire engineers. Your job is to make sure they're building against your vision, not theirs, and that the outputs are what you
want." He wasn't being gentle about it. He was being accurate by describing the same dynamic that exists for any non-technical founder working with an engineering team or any domain expert team: the function of leadership isn’t to be the most technically fluent in the room. It's to be the clearest about what you want and then let the experts
solve it with your vetting.
Armed with his shared wisdom, I did next was practical for me. I reconfigured Cashew to produce what I started calling “convenience translations”. Now every technical readout has two layers: the technical content sits on top for the record and, underneath that, what this means in plain terms—what decision is required by me, what the recommendation is, and what the risk profile of each option looks like and if any decisions would be difficult to walk back.
I'd built a version of this concept before when I was standing up new entities in countries where English wasn't the operational language and/or I wasn’t as familiar with their jurisdictional laws. My legal teams would produce the legal documents in the local language, with English translations alongside, and attorney sidebar commentary on any jurisdictional quirks that wouldn't be obvious from the text alone. I framed the technical documents that same way: don't simplify the source document but give me the context to knowledgeably use the document.
What I learned from building that translation layer isn't that I needless technical rigor, instead, the access to rigor requires legibility. You can't approve what you can't read.

The Path to AI Fluency Starts by Really Using It
So what is the point of all this?
There's a version of AI fluency spreading through HR right now that I'd describe as a house of cards with a pretty UI. Anyone can vibe code a clean dashboard, a useable chatbot or even an onboarding that runs on a reasonable script. These things can be built in an afternoon with the right prompts yet very little understanding of what's happening underneath.
They're also frequently insecure—API keys exposed; no access control; notesting protocol; and/or no documentation of what the system does under failure conditions. The people who built them don't know these things because they've never had to think about it. The AI just builds what you tell it, but little more. The failure modes may not show up if the tool isn’t widely used or consequential enough to surface then.
I know this because I was that person. The Replit app I built in the early days worked. Yet, it would not have survived any meaningful scrutiny, let alone HIPAA requirements. I didn't know enough to know what I didn't know.
Four months of structured, disciplined, documented building later: I finally know what I don't know.
The most successful HR and AI transformation leaders will be the ones who can tell you what they've built; what broke; what they approved that they later wished they hadn't; and why they made each call. They will have the kind of knowledge that doesn’t come from reading vendor decks and attending conferences. It is the kind that comes from being accountable for an outcome.
You can't lead what you haven't tested. Testing isn't vibe coding a UI, patting yourself on the back when it works, and calling it a product. Testing is running something until it fails, understanding why it failed, fixing it, and then approving the fix with a clear head.
What's Next
The second six weeks is when my agents build the thing.
I don't know what comes after that. I'm not going into this expecting aparticular outcome beyond understanding something that I currently only half-understand, which is whether the product I've been thinking about the last several years can be built by an agent team led by someone who knows the problem better than the tools.
If it can, that tells me something important about what's possible for HR leaders who are willing to do the hard, unglamorous, undocumented work of truly building and advising CEOs on what agents can do, and how to accelerate human potential with them. If it can’t, that is equally valuable.
Either answer is more useful than an opinion without a named experience.
