Full installation of the AI platform in your server room or private cloud: sizing, hardening, models, RAG over your documentation, directory integration and an offensive verification before handover.
The six phases
1
Final sizing
Using the proof of concept data: chosen model, real concurrent users and acceptable response time. Hardware follows from that, not the other way round.
2
Hardened installation
Hardened base system, minimal surface, network segmentation, encryption at rest and in transit, and secret management. The platform is treated as what it is: a critical asset with access to sensitive documentation.
3
Models and inference engine
Deployment of the selected models with a serving configuration matched to your usage pattern. Versions pinned and reproducible.
4
RAG over your documentation
Ingestion, chunking, indexing and retrieval tuning. The phase that most determines perceived quality and the one that takes the most work.
5
Identity and traceability
Directory integration, permissions by group and by corpus, and a record of queries and answers so governance can be demonstrated.
6
Security verification
Before signing off: prompt injection tests, RAG exfiltration tests and platform exposure checks. If something gives, it is fixed and retested.
Sizing orders of magnitude
No prices, because hardware moves too fast to publish them honestly. Criteria instead, which is the part that does not change.
Small team
Narrow case, few users
A mid-sized model serving a small group with intermittent use. Solved with a single professional card and a conventional server.
Department
Dozens of concurrent users
Here memory bandwidth and batch processing capacity rule. This is the tier where most deployments fall short, having bought looking only at total memory.
Organisation
Cross-cutting and critical use
Multiple nodes, load balancing, high availability and environment separation. It stops being a server and becomes a service with its own continuity plan.
The most frequent purchasing mistake is choosing by total memory. A machine can load an enormous model and serve it at a speed that is useless for several people at once. Real concurrency is the criterion, and that is why sizing comes after the proof of concept and not before.
What is delivered documented
Deployed architecture, with a network diagram and justified decisions.
Procedure for updating models and reindexing the corpus.
Permission matrix: which group reaches which documentation.
Security verification report, with what was found and what was fixed.
A usage guide for the people who will actually use it, in their language and without jargon.
Continuity and backup plan, including what to do if the machine goes down.
And afterwards
A deployment with no operation degrades. Models improve, documentation changes and answers age. Your team can run it with the documentation we hand over, or we can, using the same hour packages we already use for ongoing consulting.
Frequently asked questions
How long does a deployment take?
Mostly it depends on the state of the documentation and on identity integration. A narrow case with a prepared corpus closes in weeks; a cross-cutting deployment with several document sources and fine-grained permissions is a project of months. We bound it after the proof of concept.
Can it run in our private cloud instead of on-premise?
Yes, as long as control stays yours and the data does not leave the scope your compliance allows. What we rule out is a third party model-as-a-service, because that gives away exactly what you came to gain.
Does it integrate with our tools?
Usually yes, through its own web interface and connectors to the document sources you already have. We pin it down during sizing, because every integration adds attack surface and has to be weighed.
Who keeps the knowledge of the system?
You do. The handover documentation is written so your team can operate it without us. If you prefer that we operate it, that should be for convenience, not because we left you with no alternative.
Let us talk about your specific case
Half an hour to review what you want to solve, for how many people and over what documentation. That determines whether the short path is a proof of concept or a deployment straight away.
Phone: 686 250 244 · Email: info@jaymonsecurity.com
You may also be interested in

