← Answers

Mid-market operations

Is our data safe with an AI consultancy?

Short answer

Your data is as safe as your contract and architecture make it, and both are entirely within your control. Require four things before any access is granted: a written no-training clause covering the consultancy and every model provider in the chain, named data residency with a Data Processing Agreement where personal data is in scope, least-privilege access to redacted or sampled data rather than production, and full IP assignment of code, prompts and evaluation sets. Enterprise API tiers from the major model providers do not train on submitted data by default — the real exposure is usually an over-permissioned integration or an unlogged internal tool, not the model provider.

Last updated August 14, 2026 · Bitfumes AI consultancy

What to require in the contract

  • No-training clause covering the consultancy, any subprocessor, and every model provider in the chain — named, not implied.
  • Data residency and retention stated explicitly: where data is stored, for how long, and what is deleted at the end of the engagement.
  • A Data Processing Agreement wherever EU, UK or other regulated personal data is in scope, with subprocessors listed.
  • IP assignment on payment covering code, prompts, evaluation sets and any fine-tuned weights.
  • Named engineers, background-checked where your policy requires it, with access revoked on rotation.
  • Breach notification terms with a defined window.

What to require in the architecture

  • Least privilege — read-only, scoped to the tables and documents the use case actually needs.
  • Redacted or sampled data for development; production access only when unavoidable and only for named individuals.
  • Access through your identity provider, so revocation is one action and access is auditable.
  • Full request logging, retained by you, so you can answer what was sent and when.
  • Document permissions respected at retrieval time — a RAG system that ignores existing access controls quietly turns every document into a company-wide document.

That last point is the most commonly missed and the most expensive to discover late. It is not a model risk; it is an ordinary authorisation bug with an unusually large blast radius.

Regulated industries

Healthcare, financial services and legal work add requirements rather than changing the fundamentals: a Business Associate Agreement or sector equivalent, audit trails that survive review, retention rules that may conflict with a model provider's defaults, and evidence of output quality rather than assurances. Budget for the review cycle in the schedule — in regulated environments, security and compliance review is routinely longer than the build.

How Bitfumes handles it

We work inside your repository, your identity provider and your access controls rather than standing up a parallel environment, and we assign all IP to you. Where personal or regulated data is in scope, residency, processing and no-training terms are settled in writing before any access is granted. If a use case cannot be built safely with the access you can reasonably give, we will say so during the assessment rather than after the invoice.

Frequently asked

Will our data be used to train AI models?

Not if you contract for it. Enterprise and API tiers from the major providers do not train on submitted data by default, but the guarantee should be written into your agreement and extended to the consultancy and its subprocessors rather than assumed.

Can we run AI on our own infrastructure instead?

Yes — open-weight models running in your own environment remove the third-party question entirely, at the cost of higher infrastructure spend and some capability gap against frontier models. It is the right choice when regulation or policy makes external processing a non-starter, and overkill when it does not.

What is the biggest real data risk in an AI project?

An over-permissioned retrieval system that surfaces documents to people who were never meant to see them. It is an authorisation failure rather than an AI failure, which is exactly why it slips past reviews focused on the model.

Do we need a DPA with an AI consultancy?

Yes, wherever personal data covered by GDPR or equivalent regimes is in scope — and it should name subprocessors, including model providers, since they are processing on your behalf too.

Related answers

Next step

Want this built into your business, not just explained?

Our AI Opportunity Assessment maps where AI saves you time and money, and prices the build — $999, a written report, 7–10 days. If the answer is that AI is not worth it for you yet, we will say so in writing.