news

Meet SPOT: The AI Reviewer Built for Engineers

Drummond Carpenter’s custom-built review tool is gaining recognition as a finalist for the AEC Innovators Award.

SPOT started with a familiar challenge: our engineers were producing more work, and senior staff only had so many hours to review it.

So we built SPOT.

SPOT reviews reports, memoranda, plan sets and other deliverables using a process based on how our own senior engineers review work. It has its own email address and Teams profile, making it easy for our engineers to send something over for another set of eyes before it goes out the door.

SPOT doesn’t make the changes. It flags what it finds, ranks the issues and makes recommendations. The engineer decides what to do with them and remains responsible for the final work.

The feedback goes both ways. Engineers can agree or disagree with SPOT’s recommendations and explain why. Over time, that has helped us capture patterns in our work, from common errors to the kinds of comments that come up repeatedly with particular clients.

John Wolf recently sat down with Insights by KP to talk about why we built SPOT, how it has changed since we first put it to work and where we see it going next. Read the full interview below.




SPOT Review: The AI Colleague Drummond Carpenter Built to Beat Its Review Backlog

Originally published by Insights by KP on August 6, 2026.


At 11:30 p.m., someone at Drummond Carpenter emails a document to SPOT and asks for a review. SPOT doesn’t mind. SPOT has its own email address, its own Teams profile, and no concept of office hours. SPOT is the firm’s AI reviewer, and the junior engineers treat it like a colleague, one with no judgment attached to asking.

John Wolf built SPOT because the firm had more work than people, and junior staff were producing deliverables faster than senior engineers could review them. Worse: everyone was checking their work with whatever model they happened to like, and two people running the same document through different AI models got completely different reviews. “I checked it with AI” had stopped meaning anything.

I talked with John about how SPOT works and what it taught him. Here’s what piqued my interest:

Recommendations only, always. SPOT never changes a document. Everything comes back as track changes and comments, ranked high, medium, low, and the engineer whose name goes on the deliverable stays responsible for every line. John is direct that this framing saved them. The trust in the tool came from what it refuses to do.

The feedback loop is the product. Every review comes back with agree or disagree options, and every response feeds the harness. When people agree SPOT’s flag was right, it hunts that error harder. When they consistently disagree, it learns to let go. Hundreds of documents in, patterns surface that no human would ever compile: errors that show up in 30% of deliverables, clients who always redline the same thing. John’s story about the CEO’s phone number is the perfect illustration, and I’ll let him tell it.

Own the harness, rent the model. SPOT’s underlying model has changed about five times, and the staff never noticed. When a new model launches, John runs it against benchmark documents; if it underperforms, back to the old one. His biggest lesson from the whole build: if you own everything around the model, the model becomes the only thing you rent. His long-term plan takes that idea somewhere most firms haven’t even considered.

The full conversation is below, lightly edited for length and clarity.

Drummond Carpenter is a finalist for the AEC Innovators Award. You can vote for them and learn more about the award at kpreddy.co/aec-innovators.

With more work than people, engineering and environmental consulting firm Drummond Carpenter faced a familiar bottleneck: junior staff producing deliverables faster than senior engineers could review them. Their answer was SPOT — a custom-built AI agent with its own email address and Teams profile that reviews documents the way the firm’s senior engineers would, before anything goes out the door. We spoke with John Wolf about why consistency drove the build, how a feedback loop turned Spot into an institutional knowledge engine, and why the experience changed how much he values owning the layer around the model. The conversation has been edited for length and clarity.

Let’s talk about SPOT. Where did the idea come from?

We have more work than people. We put out a lot of content — reports, memoranda, plan sets, any form of deliverable — and a lot of engineering work is kind of a black box, similar to AI: they give you something to work on, you go work on it, and the deliverable is what you produce. The bottleneck for us was making sure deliverables were reviewed correctly. The engineers have their own work; they can’t constantly review stuff. And because we hired so many young engineers, they were either nervous to ask for a review or, worst case, it wouldn’t get reviewed and would go to a client.

So we built a custom agent harness. It’s built into our company — it has a profile, it acts kind of like another employee. It’s an intermediary step where someone can email it and get a review the way our senior engineers would review it, so junior people can get things reviewed the way our company believes they should be reviewed before they go out the door. It’s really helped our people maintain consistent reviews. That’s a big deal for us: you have to check stuff, and you have to check stuff in a very specific way.

How did you go about building something you felt was confidently doing that check-and-balance process?

We were using regular ChatGPT, and we noticed it was pretty inconsistent. Two people would send the same document and get slightly different reviews back, which was scary. And people constantly switched models — some were using Claude, some ChatGPT, some Gemini. It’s like, “Oh yeah, I checked it with AI.” Well, you checked it with Gemini and I checked it with ChatGPT, and they’re totally different.

So we really wanted something that would review everything consistently. We sat down and mapped out exactly how I would review something, and had that process checked by a bunch of engineers. Little things like: if you see a callout for a construction detail, you flip to the back of the set and make sure the detail matches. We mapped out that entire process.

How long did it take to build?

It’s still under development, in a sense. One of the cool things we’ve figured out is that when you submit something, the email you get back has “agree” or “disagree” options, so the person getting the review can push back. We had one where SPOT said a 30% contingency at that point in the design seemed high, and the engineer said, “No, I disagree, for this reason and this reason.” Those things are constantly fed back into the harness. But the original product rolled out about two or three months after we started developing it.

What has the reception been from users?

Extremely positive. Younger people especially feel there’s no barrier to asking for a review — it’s essentially free for them, and they can do it any time of day. I get little notifications and see people asking for reviews at 11:30 p.m., which is awesome.

It’s also helped that it’s caught really valuable things — even silly stuff, like the header date being wrong, or using the client’s name inconsistently throughout. It’s not just engineering things. People are constantly in our company Teams chat saying, “Look at this cool thing it found — I would have never caught this.”

What was the decision behind giving it an email address and a Teams profile, rather than a chatbot or standalone tool?

We wanted it as embedded in the workflow as possible. Everybody’s in Outlook all day, everyone’s in Teams all day, and that’s typically how you ask for a review — you email it to somebody or send a SharePoint link. We considered a small browser-based agent, but we felt that putting it where all the other co-workers are is helpful. That way you can CC people, so your supervisor knows you asked for a review, can see the findings, and can see how long your list is.

The other thing is that we’re very security conscious. Building it in our Microsoft tenant — where Outlook and Teams are already behind our wall — prevented a lot of the IT concerns of an external product where people had to leave our company wall to work.

Have you done any time or impact studies on the ROI?

We think it’s about five hours per document, just from the standpoint of someone actually sitting down, reviewing, and going back and forth. It lays the findings out in an Excel sheet if you want, so you can share it with a group, and it’s interactive — people can say “I did this one” or assign items to each other. And that’s just the review itself. If somebody emails me something to review, it might take a day and a half or two days before I even open it — that part is harder to quantify. But at our billing rates, five hours per document alone is substantial.

Where’s the line today between what you trust SPOT to flag and what still needs human judgment?

That’s a really good question, and I don’t have a good answer. What I can say is we’re very specific that it doesn’t change anything. If you send it a Word document, it marks it up in track changes and adds comments, but it doesn’t make the changes. Everything is just a recommendation. You never get back a document with changes you don’t know about.

It ranks findings high, medium, and low in severity — an incorrect engineering calculation is high severity, but it’s still just a recommendation. We’re explicit that the human is responsible for going line by line. It’ll say, “You didn’t use an Oxford comma,” and if you want to ignore that one, that’s up to you. At the end of the day, it’s not SPOT sending a deliverable — the engineer putting their name on it is responsible. It’s a tool to augment you, not replace you. Framing everything as recommendations has saved us a lot.

Is there anything your submission doesn’t capture — a learning that changed your opinion about AI or how you work?

The most eye-opening thing, for me personally, is how much more I now weight ownership. Our staff don’t even really know this, but the model behind SPOT has changed about five times. We’re completely model agnostic. We have benchmarks — documents reviewed by both a human and the AI, where we understand what should get flagged. When a new model comes out, like when Fable came out, we plugged it in, ran it through, it didn’t do as good a job, and we went back to our old model. We can do that because we own everything around the model layer. I originally built it because I wanted consistency, but that turned into the understanding that if we build the agent harness, we control everything except whatever model we decide to rent.

The other thing is the feedback loop. Capturing it has built up a massive knowledge base of what usually gets flagged. If it flags something and a lot of people agree, it starts understanding that it’s a common error and looks for it more. If people consistently disagree, that becomes institutional knowledge we feed back into the harness, and it de-risks it. That knowledge layer has gotten really valuable. When you send this thing into the wild with hundreds of documents flying through it, it starts to surface patterns — this thing shows up in 30% of deliverables, or a lot of submissions with this client have this incorrect detail — things you’d never think you could capture.

Do you have to instruct employees on how to leave comments so the feedback is most useful?

There is a system. You get the Excel back, give your disposition — did I implement it, yes or no — and why or why not. It’s pretty straightforward. But we absolutely had to coach people on why providing feedback is so important. For example, our CEO’s phone number is on all of our plan sets. If you sent something out of Michigan, SPOT would flag: “You have a Michigan address on the title block but a Florida area code in the phone number — I think that’s weird.” People would just ignore it rather than give feedback. But when people did give the feedback, things like that stopped showing up. Then they understood — it’s like helping an intern. If you provide feedback, they actually get better. Once people saw the flywheel turning and the reviews getting way better, it was much easier for us.

What’s next? You’ve captured a lot of common wisdom — is there a plan to go beyond what you’ve done today?

We have a lot of plans, and a lot of what we’ve built has adapted to this mission statement we now have: if you’re going to do something with AI, you should own the understanding of how it worked — what worked and what didn’t. So with SPOT and other processes, when someone interacts with an agent, we’re capturing the overall interaction. Currently that’s mostly to improve the harnesses. We’ve gotten really good at harness engineering and context engineering, so the stuff our employees interact with is almost always custom-built — which has been cool for getting us off the ChatGPT-cloud dependencies.

Long term, my goal is to distill local models so we can be completely independent. Instead of relying on frontier APIs, I’d like to develop in-house local models that might not give you a good cocktail recipe, but are really, really dialed in on how we do a plan set review. We’re going to have hundreds of thousands of input-output pairs, graded by actual engineers — this is a good comment, this is a bad comment. Once we build that up, I’d like to start distilling our own models.

But how we surface the intelligence could use work. We capture so much — things like “this client typically takes six weeks to approve a construction permit” or “this client always redlines that one thing.” How and when we surface that is something we’re still figuring out. I think there’s a larger business opportunity in that context and institutional understanding, but truthfully, I don’t know what it is yet.

Has anything surprised you? You’re a pretty small firm — any impediments or lessons that would be useful for others?

The biggest impediment has been how inconsistent AI use can be across the company. People have such different workflows and preferences that what you get back drastically differs. Our biggest hurdle has been creating a consistent workflow without diminishing or discouraging exploration — we’re very innovative, and we’re lucky our leadership is supportive of using AI however you want. But that leads to inconsistent use. Two people will draft something and it’s written completely differently, and as a company you want everything to seem like it came from the same place.

That’s another reason we started building things ourselves — so we’d have a little more control over the actual prompts on the back end, the guardrails, all of that, without discouraging people from putting everything through AI, because we do believe in that. Walking the line between enforcement and encouragement has been difficult. There’s a responsibility for the company to maintain its identity and standards, and injecting that into a process that’s moving so quickly has been tough.

That’s a pro and a con of being small — everyone can jump in and solve their own problems, but you can’t have much governance infrastructure.

You’re looking at it — this is me. And I have this theory that the people who use AI the most end up doing more work. It’s much higher-level, more cerebral work, where you’re making big-scale decisions. I remember years ago being able to put on a podcast and just do my work. That time is gone now. It often feels like I have multiple jobs, which is cool overall.