Hi Adopter,
If your team handles private information, you’ve probably asked whether AI can help without sending the work to another company’s model. Most conversations jump straight to cost. Honestly, that misses the more important choice.
Who controls the model, where the information goes, what gets logged, and when a person takes over? I call that intelligence independence. You don’t need to build a frontier model. You need a real choice in how your business uses one.
A law firm’s very small AI job
In an anonymized case study from Nexus Technology Consulting, a Sydney boutique law firm had a specific problem: a partner returning from leave faced a mountain of email and calendar items before they could get back to client work.
Nexus says it put an open-weight, roughly 8-billion-parameter model on the firm’s own RTX 4090 graphics card. The system pulls messages from the firm’s Google Workspace account, writes a one-line summary for each email, and puts a daily rollup into a local spreadsheet. The partner reads the rollup, then opens the original messages before acting on anything important.
The vendor reports that one longer-leave catch-up fell from about two hours to 25 minutes, and that the system summarizes about 210 emails per hour. Those are vendor-reported figures from an unnamed client, not independently measured ROI. The exact model checkpoint, total build and running cost, and error rate aren’t public.
Still, the design is worth studying. The model was given a small, repeatable reading job. It wasn’t asked to decide whether a deadline had been met, advise a client, or send a reply.
And “local” needs a little honesty. The AI inference runs on the firm’s machine, according to Nexus. The email still lives in Google Workspace. Local AI removes one outside AI processor from this workflow; it doesn’t make the whole information chain independent or automatically secure.
The choice it creates
Madgett Law describes a different use of local AI. Its published approach gives clients a choice in their engagement letter: no AI, in-house AI only, or a hybrid that can use contracted outside providers. The firm says the in-house-only choice must be enforced by the routing software, not left to someone’s memory. A lawyer still checks the work and verifies citations in every tier.
That’s a firm-authored policy, not proof of a fully local deployment or a measured saving. But it points to the real prize: a business can make a promise about how work will be handled and have a way to honor it.
Below, I’ve put these cases beside North Dakota’s local bill-summary pilot and built a decision test you can use on one workflow. The full case-study PDF includes the stack evidence, the missing numbers, a security checklist, and a worksheet for deciding whether local, cloud, or hybrid is right for your team.
Want help choosing the right path?
AI Adopters Coaching starts with one bottleneck in your business. If local AI fits, we can build a useful workflow and teach you or someone on your team to run and check it. The intake is free, with no commitment. You don’t need to buy hardware to start.
What three legal-work cases actually show
Nexus is the most concrete small-firm hardware story, but it is also the least independently checkable. We have a vendor’s description of a 38-person firm, an RTX 4090, an “LFM2-8B class” model, a local spreadsheet, and a partner reviewing original messages. We don’t have the firm’s name, test set, security audit, cost, or before-and-after task logs. I wouldn’t build a return-on-investment slide from it.
North Dakota is a useful counterweight. Meta’s account of the Legislative Council names Llama 3.2 1B Instruct, fine-tuning with Unsloth, a retrieval step that brings in relevant bill text, and Ollama running on local hardware. The system generated three draft summaries per bill for legal review during the 2025 session. Meta calls its 15–25% time saving and 25 legal hours per session estimates. The 2027 use is a plan, not an observed result. It is a public legal team, not a law firm, and its exact machine specifications and operating costs aren’t disclosed.
Madgett supplies something neither case gives us: a reason to own the option. A client can choose a handling tier because the firm says its routing follows that choice. The published article gives hardware and model-size guidance, but no measured workload, deployed checkpoint, or ROI for the firm. Treat the tier design as a practitioner example, not a legal standard.
Put together, they show three forms of control: a bounded task on owned hardware, a locally served model adapted to a legal task, and a business rule about when outside AI is allowed. None proves that every job should run locally.
What you need to own
If you want this kind of independence, start with four questions. You don’t need to buy a server to ask them, but someone must own the answers.
What information can the model see? Trace the data from its original system to the model, the output file, logs, backups, and any human inbox. In the Nexus story, “local AI” still starts with Google Workspace.
Who controls the model version? Name the exact checkpoint and its license. Keep a tested copy you can roll back to. An open-weight model can still carry use restrictions.
What is the model allowed to do? Summarizing for review is different from filing, sending, advising, or making a decision. Put the handoff to a person at the point where an error has consequences.
Can you switch paths? If a provider changes price or terms, can you move the same task to another approved model? If the local model isn’t good enough, can you route a carefully chosen subset to a contracted provider, or keep it human-only?
That last question is why I don’t think this is only a cost story. A cheap API can be the right choice. So can a managed cloud service with strong controls. The question is whether you’re choosing it, or whether your business has quietly lost every other workable option.
A test small enough to finish
Pick one task that produces a draft, not a final decision. Email catch-up, document tagging, or a first-pass summary are reasonable candidates. Use 20 real, permission-cleared examples as an initial screen, not as proof of safety. Include ordinary cases, awkward cases, and cases where the right output is “I can’t tell.” Keep the original task and source documents available to a reviewer.
Before testing a model, record how long the current process takes and what an acceptable result looks like. Then compare three paths on the same examples: the current human process, a suitable managed service if your rules permit it, and one local model that fits the machine you actually have. Don’t buy a graphics card because a case study had one.
Have a qualified reviewer score correctness, missing information, source traceability, and whether the system knew when to stop. Add the person’s checking time, setup time, maintenance, electricity, and hardware to the local cost. For a cloud path, include usage fees, contracts, administration, and the cost of keeping sensitive material out of it. This is a business comparison, not a benchmark contest.
I would stop the pilot immediately for an access leak or a wrong answer in a predefined high-consequence case. Otherwise, decide at the end whether the result is good enough to continue, needs a narrower task, or should stay with the current method. A local model that makes everyone double-check every line hasn’t freed anyone.
Download Today’s AI Case Study 👇
The full case study and worksheet give you the evidence table, architecture and control map, a cost sheet, and a one-page continue/change/stop decision. The point isn’t to convince you to run everything in your office. It’s to help you know which choices you’re making, and which ones a vendor has made for you.
What is one piece of work in your business that you would want the freedom to run on your own terms?
Adapt and Create, Kamil






