Every software demo is built to show you the answer arriving. Very few are built to show you where your question went. If your files include customer records, payroll, patient charts, or anything a regulator reads, the second thing matters more than the first.
The trip your question takes
When an employee pastes three paragraphs about a named customer into an online AI tool, a specific sequence happens. The text leaves your building and lands on the vendor's servers. It is logged. It is retained on the vendor's schedule, not yours. Depending on the account tier, it may be reviewed under the vendor's terms and used to train future models. It may pass through sub-processors — companies you have never heard of and did not choose. Nothing improper occurred at any step; everything printed in the terms of service happened exactly as printed. But your customer's story is now a record on infrastructure you do not control, governed by terms you did not write, subject to legal process aimed at someone else.
Training is a tier setting, not a promise
Take the best-documented example. OpenAI's published policy is that conversations on personal free and paid accounts are used for model training by default, while business and API tiers are opted out by default. The setting exists and it is honest. The question is which tier your people are actually on. The 2025 Enterprise AI and SaaS Data Security Report from the security firm LayerX found that 82 percent of workplace pastes into AI tools came from unmanaged personal accounts — not the account the company vetted, the one the employee already had. Samsung learned this the direct way in 2023: within weeks of permitting ChatGPT internally, engineers had pasted source code and a meeting transcript into it on three occasions, and the company banned the tools outright. The gap between the account the business bought and the account the staff use at nine at night is where data walks out.
Retention can outlive your delete button
In the New York Times copyright litigation against OpenAI, a federal magistrate ordered in May 2025 that output logs be preserved — including conversations users had already deleted — and the company was later ordered to produce twenty million de-identified chat logs to the plaintiffs. The preservation obligation was substantially lifted that October, but the structural point does not lift. Your retention schedule, the one your attorney or compliance officer wrote, is subordinate to litigation you are not a party to, for exactly as long as your data sits on a vendor's infrastructure. No contract you sign with a vendor prevents a court order against that vendor.
Six questions to ask before any demo
The demo can wait. These cannot. First: where, physically, is our text stored when we type it in? Second: is our data used for training, and under exactly which tier of your terms — can you show us the setting? Third: how long is it retained, and who can order it retained longer? Fourth: which sub-processors can see it? Fifth: when we leave, what gets deleted, and what do we receive that proves it? Sixth: will you put each of those answers in the contract?
A vendor who answers in plain sentences has earned another meeting. A vendor who answers with a brochure has also answered, just not the question. Ask us the same six. A shop that builds private systems should welcome them, and you should be suspicious of one that does not.
What running it on your own hardware changes
The architecture we build puts the model on a machine you control — in your building or on your rack. Your documents are indexed there. The question, the retrieval, and the answer never cross your network boundary. That removes an entire category of exposure: there is no third-party processor holding your records, no vendor whose subpoena becomes your discovery problem, no tier setting to audit every quarter.
On compliance, plainly: this keeps records on infrastructure you control, which is the starting point for a compliance conversation with your own counsel or compliance officer — not a substitute for one. No certification is claimed here, and none could honestly be. For the medical trades in particular, HHS does not certify any product as HIPAA compliant, so a vendor claiming a certified product is telling you something false by definition.
What it does not change
A machine in your own closet does not protect you from yourself. The laptop that walks out of the truck, the backup nobody encrypted, the technician pasting into a personal app on his own phone — local hardware touches none of that. It also creates work: patching, updates, and a locked door become your responsibility, or the responsibility of whoever maintains the system for you. The honest summary is that on-premise architecture removes the third-party category of risk and leaves every other category where it always was — with you. Anyone presenting local hardware as making the data question disappear is doing marketing, not architecture.
Where to start
Ask your own team — anonymously, if that gets truer answers — which AI tools they already use for work tasks. The result tells you whether your data question is about some future vendor or about last Tuesday. Then put the six questions in writing and send them ahead of the next demo you sit through. Anyone's demo, including ours.