Sensitive data travels inside the prompt
Every call to an external model can move regulated information outside the perimeter, with no log, no control, and no way to prove otherwise.
We design, deploy and govern AI capabilities on your own infrastructure. Full traceability, verifiable data sovereignty, and evidence ready for an audit.
Three gaps show up in almost every diagnostic we run, regardless of sector or size.
Every call to an external model can move regulated information outside the perimeter, with no log, no control, and no way to prove otherwise.
Use cases get tested, but without a baseline and an agreed metric there is no basis for deciding what to scale and what to drop.
Data protection rules and ISO 42001 require inventories, logs and audit trails. A signed policy is worth nothing without operational evidence behind it.
Adoption with results
Every initiative starts with a baseline and ends with an agreed metric. From diagnostic to inference cost control.
360° Check-up covering maturity, risks and roadmap. Claude, ChatGPT and Gemini Enterprise adoption, with licensing, policies and training. AI FinOps across AWS, GCP, HWC and Azure, and across APIs. Comparison between proprietary and open-weight models.
Data ready and protected
AI never touches production systems. We build prepared, controlled data layers so the model consumes only what it should.
Controlled replicas and vectorization. Document conversion into AI-readable formats. Data leakage audits across LLM traffic.
Sovereign infrastructure
We size the compute and deploy open-weight models on the client's own AI infrastructure. Inference happens inside the perimeter.
Sizing and power consumption for NVIDIA GPU, AMD GPU, Huawei GPU, Google TPU, or Apple M5 Max and Ultra. Assisted deployment of Qwen, Gemma and DeepSeek. In-house lab for sovereignty testing.
Our own products, in operation. We do not sell what we cannot demonstrate.
Control gateway, part of AI Vault Management
A gateway between your applications and any model. It detects sensitive data before it crosses the perimeter and decides what to do with it.
The request never leaves. Used when context or regulation make any substitution unviable.
The value is replaced irreversibly. The data is anonymized and the model keeps the task context.
The value is swapped for a reversible token and restored in the response. The data is still personal and still under control.
Every decision is logged. The evidence is auditable.
Consumption and emissions measurement.
Every inference consumes energy. Huella IA measures the consumption and emissions of your AI workloads across AWS, GCP, HWC, Azure and APIs, and reports them alongside cost.
Model consumption, part of AI Value Management
A prepaid wallet of AI tokens in local currency. The client tops up, picks a model and consumes it from a web chat or through the API, deducting per token used.
You pay for what you use. No lock-in contract and no fixed per-seat fee.
Top-ups happen in local currency, with the payment methods the organization already uses.
What you load stays available until it is used, with no expiry and no loss through inactivity.
We train models on our own hardware and publish them openly on Hugging Face, with their evaluation and provenance. It is our contribution to the community: fine-tuned models built by AVIM, available for anyone to review, use and challenge.
Gemma 4 adapted to Latin American Spanish. Contains the full evaluation, provenance and known limitations.
Open model cardThe same adaptation with weights already merged, to run with MLX without mounting the adapter separately.
Open model cardEvery new model ships with the same card: what was trained, on what data, what improved and what did not.
See the organization on Hugging FaceIn operation
128 GB of unified memory. This is where our first public model was trained, in a single run.
In operation
Extends the lab to test GPU deployments before a client buys any hardware.
In operation
Agentic capability prototyping to validate workloads on AI-based systems.
In operation
Traces, cost and latency for every call, which is what later supports FinOps and audit evidence.
In operation
Routing across models to compare quality and real cost before recommending one.
In operation
Qwen, Gemma and DeepSeek, deployed and measured on the same hardware the client will use.
Banking secrecy, regulator requirements, and models that must be explainable.
Clinical records and sensitive data under heightened consent requirements.
Public procurement, traceability of administrative acts, and sovereignty over citizen data.
Critical operations, industrial property, and sites with limited connectivity.
The diagnostic is the entry point. No infrastructure gets sized before we know what for.
Maturity, use case inventory, risk map and prioritized roadmap.
We run your critical case on an open model in controlled infrastructure, using test data.
Infrastructure, data layer and controls, with agreed metrics from day one.
Monitoring, human oversight and control of the real inference cost.
The diagnostic delivers a maturity assessment, a risk map and a prioritized roadmap. It is where everything else starts.