(2026-06-29) Jones Run This4question Test Before You Let Any Ai Into Your Files Your Slack Or Your Phone

Nate B Jones: Run this 4-question test before you let any [AI] into your files, your Slack.com, or your phone. I think the news is pointing at one shift, and it is not the one the headlines are arguing about. The AI race is becoming a race to context. The question is no longer only “who has the smartest model?” It is who can put usable intelligence closest to the work in front of you.

That changes what matters for the rest of us. If the newest model is gated, even for a few weeks, the tools already in your hands get more valuable, and the cheaper open models pressing from below get room to look competitive in public. The advantage stops being the model and starts being placement.

Where your context lives is going to matter more than which model is half a point smarter this week.

I built a prompt for that: paste it into your AI to talk it through, or point an agent that can already see your tools straight at it. Either way it runs the four questions and the exit test against the tools you actually use, then hands back which one is worth more of your context, where you’re getting hard to extract, and the one move worth making this week.

The empty chat box is a tax

The model can be brilliant and still fail if it doesn’t know what is going on.

So the most important AI products now are trying to do something deeper than improve the model. They are trying to move the model closer to the situation.

Apple is doing that with the phone. Anthropic is doing it with Slack. OpenAI is doing it with Codex and the file-shaped work surface. Chinese open models like GLM-5.2 are forcing the question from the other direction: if cheap intelligence is good enough for more tasks, then the surrounding system matters more.

Siri is the personal context bet

I think the story is that Apple is trying to make Siri useful by putting it closer to the context of your life.

If I ask a generic chatbot, “When is my mom landing?” the hard part is not language. The hard part is figuring out what is actually happening. Did she text me a flight number? Is the pickup in my calendar? Is there an email confirmation? Did my brother say he might go instead?

That is Apple’s argument, and honestly it is a pretty good one.

If Siri is reading your messages, finding your emails, interpreting your photos, watching what is on your screen, and acting across apps, the custody question becomes central. Where does the context go? Who can see it? What leaves the device? What is remembered? What can be audited or revoked?

Claude Tag is the team context bet

Move from the phone to the company and the same problem gets messier fast.

Anthropic’s Claude Tag launch sounds simple on the surface. Claude starts in Slack

**Companies like to pretend that work lives in systems of record (system of record). Some of it does. A lot of it lives in the messy middle. The customer complaint is in one channel. The pricing caveat is in another. The real decision happened in a call. The follow-up was in a thread. The official doc is stale. The current version is in someone’s drive. The reason the plan changed was obvious to everyone present and invisible to everyone else.
Without that context, even a very good model is still waiting for somebody to brief it.
Claude Tag is Anthropic saying: stop bringing every messy fragment to Claude. Let Claude get close enough to the work that it can build context over time.

The more useful Claude becomes in Slack, the more it needs access to information companies are bad at governing: engineering decisions, customer tickets, private channels, people issues, legal caveats, pricing debates, sales details, and casual comments that were never meant to become corporate memory.

So when Anthropic spends time on separate Claude identities, scoped memories, channel permissions, admin controls, spend limits, and logs, I would not skip that as enterprise boilerplate. That is the trust architecture. Without it, “AI teammate” is just a friendly name for a context leak.

The company that wins team context will do more than answer questions. It will become part of how work remembers itself.

Codex shows what happens when context can execute

If you don’t code, Codex is easy to wave away.

It looks like a developer tool

Software is the cleanest test case for context-aware AI because the work surface is unusually inspectable.

The habit here is not “developers will type less code.” The habit is delegation inside a work surface. Give the agent the environment. Let it inspect the actual files. Let it try the task. Make it leave an artifact a human can review.

Codex matters here because it shows what changes when an AI is placed inside the work instead of being briefed about it from the outside.

Most knowledge work doesn’t yet have that kind of surface.
A memo can sound right and miss the actual decision. A customer analysis can be neat and ignore the one message that mattered

In code, a test can fail. A diff can be inspected. A build can break. The work can push back a little.
That doesn’t make software easy. It makes the agent loop more legible.

For everyone else, the point is not “make all work look like code.” The point is simpler: if you want AI to do serious work, it needs context, tools, boundaries, and receipts. The assistant has to know enough to act, have permission to act, and leave behind evidence that a human can check. (legibility)

Open models make this a deployment race

The context race would matter anyway. GPT-5.6 makes it immediate

If the newest model is available only to approved partners for a while, everyone else has to get more utility from the models already in their hands.

Second, open and cheaper models get more time to look competitive in public

GLM-5.2 belongs in the story not because every company should suddenly switch serious work to a Chinese model. That is too simple. It belongs because Chinese open-weight models are making raw intelligence feel less scarce at the same moment U.S. frontier access is getting more conditional.

The last mile is still hard.

An open model does not magically know your customer’s history, your repo, your Slack decisions, your legal boundaries, your procurement rules, or your definition of done. None of that ships with the weights. It does not solve permissioning, logging, audit, data residency, evaluations, review queues, or trust. A downloadable model can win the model call and still lose the work system.

Open models are well positioned when companies want control, portability, and lower cost. Closed platforms are well positioned when they already sit where the context lives. Apple has the phone. Anthropic is trying to get into Slack. OpenAI has Codex and the file-shaped work surface.

Before asking which AI tool to use, ask four questions about the work:

  • What can it see?
  • What can it do?
  • What does it remember?
  • How do I check it?

if you are waiting for the next model to solve your work, you may miss what Apple, Anthropic, and OpenAI are doing right now: getting closer to the places where your work and life already happen.
I would pay attention here.

The same context that makes these tools useful is what makes them expensive to leave

Who holds the context? Can you export it? Can you inspect it? Can you revoke it? Can you route it to another model? Can you keep sensitive context local

This week, we're going deep on shared context — the practical kind, not the theoretical kind. I'll walk you through the Open Brain, Skills and Engine ecosystem and show you exactly how the three working together earn their keep: cutting the time you lose re-explaining yourself to a model that forgot everything you told it yesterday, and surfacing the spots where AI can actually move something in your life instead of just sounding like it might.


Edited:    |       |    Search Twitter for discussion