Brainsystems® in-house AI in practice

In-house AI for companies – local on your own hardware

With a Brainsystems® in-house AI system, you can accelerate a wide range of time-consuming work processes: summarize documents, search company knowledge, pre-screen contracts, analyze tenders, evaluate technical documentation, and prepare decisions.
Processing takes place on your own hardware within your infrastructure. Internal documents, approved data sources, queries, and responses remain within the company – without external cloud processing and without the additional security, data protection, and compliance issues that can arise when confidential content is processed by third-party AI services.

Local AI server with software stack and user interface

Practical examples of how local AI can support you:




Why in-house AI?

Local. Secure. Independent.

In-house AI: data sovereignty and digital autonomy

In-house AI creates a dedicated AI infrastructure directly within the company. Models, queries, documents, and company data are processed in an environment you control. This makes AI viable for long-term use – for confidential information, core business processes, and the development of your own digital expertise.

Company data stays internal

Documents, queries, and generated responses are processed within your own infrastructure. Companies decide where data is stored, which systems can access information, and which users are allowed to use which content.

Your own security and access concepts

Brainsystems® can be integrated into existing data protection, security, and compliance structures – with defined data sources, role-based permissions, separate knowledge areas, and internal approval processes.

AI performance without token dependency

The available computing power can be used across the company. Frequent queries, extensive documents, and growing usage are processed within the available system capacity – without usage-based cloud or token subscriptions.

Operation within your own network

Local AI operates directly within the company environment. Employees and connected systems access an internally provided platform. Availability, maintenance windows, and priorities remain under your own control.

Company-specific pipelines and workflows

In-house AI can be integrated into your own processes step by step: from document search and internal knowledge queries to database, ERP, CRM, or ticketing system connections. This creates your own AI workflows that fit existing processes rather than being dictated by an external platform.

Example illustration: Local AI system with knowledge base, internal documents, and know-how - without the cloud.


Brainsystems® Product Family

Choose the right local AI for your business

Brainsystems® systems are powerful on-premise AI platforms for businesses – from compact local AI to large-memory platforms for extensive knowledge bases and demanding AI workloads.

On-Prem AI Brainsystems 64G

Brainsystems® 64G

Typical Use Team, office, department, development
64GB ECC Pro Memory for local AI, RAG, knowledge search and document analysis
1TB Local Storage High-Speed NVMe Storage
AI Workloads Ideal for local assistants, document search, knowledge queries and clearly defined AI tasks
Multi-GPU Ready scalable to up to 4 GPU accelerators for more parallel AI requests
Configure now
On-Prem AI Brainsystems 128G

Brainsystems® 128G

Typical Use Team, department, growing AI use
128GB ECC Pro Memory large reserves for knowledge bases, RAG and parallel AI services
2TB Local Storage High-Speed NVMe Storage
AI Workloads For extensive knowledge systems, larger document collections and multiple AI services running in parallel
Multi-GPU Ready scalable to up to 4 GPU accelerators for more parallel AI requests
Configure now
On-Prem AI Brainsystems 256G

Brainsystems® 256G

Typical Use large teams, central in-house AI
256GB ECC Pro Memory Large-memory platform with maximum reserves for local AI
4TB Local Storage High-Speed NVMe Storage
AI Workloads For very large knowledge bases, extensive data processing and demanding parallel AI workflows
Multi-GPU Ready scalable to up to 4 GPU accelerators for more parallel AI requests
Configure now

Brainsystems® Platform

Built for Local AI.

All Brainsystems® systems are based on a professional, modular server platform with dedicated AI acceleration and extensive expansion reserves.

28-Core CPU AMD EPYC Pro Chip for parallel data processing, RAG, embeddings and AI services
High-Speed AI Dedicated GPU acceleration with high AI Memory Bandwidth
Dual 10GbE Two integrated 10-Gigabit Ethernet interfaces for enterprise networks and High-Speed Storage
Integrated Remote Management Remote access for system monitoring, diagnostics and administration independent of the operating system
Modular by Design Tower system for office use or 19" rack – with extensive expansion options for future requirements
StorageCube Ready Direct 10GbE connection to optional StorageCube® for large local data and knowledge bases

Brainsystems® is a local AI hardware platform. The customer is responsible for the selection, license review, installation, configuration, and use of individual AI models. AI models mentioned are used solely as a technical reference for memory requirements, runtime behavior, and system sizing. They are not part of the standard scope of delivery and do not constitute a recommendation, approval, or assurance for production use.

System benchmarks · concurrent AI queries

More GPUs – More users
More GPUs do not reduce the response time of individual AI queries.
They enable more concurrent AI users.


The appropriate Brainsystems® configuration level depends not only on the number of employees, but on how many AI tasks should start smoothly at the same time. The benchmark shows how 1, 2, and 4 GPUs affect response start under parallel use.

User experience is determined not only by the raw compute performance of a single query. In practice, what matters most is how the system behaves when several people use the in-house AI at nearly the same time: summarizing documents, comparing content, querying internal information, or starting longer analyses.

Assuming a desired response start of approximately 30 seconds, the benchmark gives the following picture:
5 parallel AI queries with 1 GPU -> 29,5s (to P95 response start)
10 parallel AI queries with 2 GPUs -> 31,7s (to P95 response start)
20 parallel AI queries with 4 GPUs -> 35,0s (to P95 response start)


It is important to distinguish between concurrently active AI queries and AI users. Not every user submits a query at exactly the same moment. The actual number of people who can work with a system can therefore be significantly higher. It depends on the type of tasks, frequency of use, document length, model size, and typical peak loads within the company.

In this context, more GPUs do not simply mean “more speed” for a single query. The main effect is additional parallel capacity: multiple model instances can work simultaneously, queues become shorter, and visible response start remains more controllable as usage increases. This is precisely what matters for day-to-day user experience.


Response start with 1, 2, and 4 GPUs

The chart shows the P95 response start for multiple concurrently active AI queries. This makes it clear when additional GPUs stabilize the user experience and keep parallel use more responsive.

X1 · 1 GPU X2 · 2 GPUs X4 · 4 GPUs Reference point

The values shown are benchmark and reference values from a defined test configuration. The reference points of 5, 10, and 20 concurrent AI queries were interpolated between adjacent benchmark levels. In practice, the model, context length, document size, runtime environment, and usage behavior can affect the results.

AI model sizes and VRAM

Response quality

Model size matters. Context is crucial.


The available GPU memory helps determine which AI models can be run locally. In simplified terms: The larger a model is, the better it can handle complex questions, long-range relationships, and tasks with many conditions. For many business tasks, however, a medium-sized model that can run with 32 GB of VRAM is already sufficient.

In practice, however, model size is only one part of response quality. The right context is often more important: internal documents, manuals, policies, contracts, technical documentation, project data, or approved knowledge areas. This company context is what turns a general-purpose language model into a useful assistant for specific internal tasks.

Documents Knowledge areas Databases ERP CRM Ticketing systems Project data internal interfaces

A Brainsystems® in-house AI system can be connected to internal data sources with specifically controlled read-only access. This gives the AI access to relevant company information without confidential content being sent to external cloud services for processing. The data sources remain controllable: read-only, isolated, role-based, and tailored to the respective application.

Larger models become particularly relevant when very complex analyses, many conditions, long documents, or especially large contexts need to be processed. The choice of models remains flexible: multiple models can be used in parallel, switched, and tested. The existing internal documents, data sources, and company context remain available.

Named AI models, model names, product names, and trademarks belong to their respective rights holders. They are named solely for the technical classification of model size, memory requirements, and local executability under specific test and configuration conditions. Their mention does not constitute scope of delivery, a recommendation, certification, ranking, or quality comparison of individual AI models, manufacturers, or providers. AI models are not part of the standard scope of delivery of Brainsystems®. Availability, license terms, and usage rights for individual models must be reviewed separately before use within the respective company.

System benchmarks · parallel AI usage

Multi-GPU scaling with concurrent AI queries

The tabs show only measured hardware, response-time, and scaling values within the respective reference model. They do not constitute a quality comparison, ranking, or recommendation of individual AI models. The results show the behavior of the local system platform under defined test conditions. DeepSeek R1 and GPT-OSS are not included in the cross-model mean without Thinking because of their reasoning output.

Qwen 3 30B

Q4_K_M · approx. 19,2 GB VRAM per model instance

With 32 concurrently active queries, two GPUs achieved 1,99 times and four GPUs 3,97 times the total throughput of a single GPU.

Time to response start during parallel use

P95 time to the first visible response token · 1 to 32 concurrently active queries

The start time indicates the first visible response token.

Total throughput at 32 active queries

1 GPU
1,00×
2 GPUs
1,99×
4 GPUs
3,97×

Interpretation

With four concurrently active queries, the P95 response start was 4,28 seconds with X1, 2,49 seconds with X2, and 0,70 seconds with X4.

Under the highest measured load, X4 achieved almost four times the total throughput of a single GPU.

Qwen 3 30B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries 1 GPU 2 GPUs 4 GPUs
P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s
1 0,71 s1,66 s153,60 0,70 s1,64 s155,25 0,74 s1,69 s152,77
2 1,64 s2,58 s193,24 0,70 s1,64 s308,02 0,70 s1,65 s308,85
4 4,28 s5,24 s194,06 2,49 s3,43 s385,64 0,70 s1,64 s616,87
8 9,42 s10,37 s198,19 5,24 s6,19 s388,95 3,07 s4,02 s765,50
16 19,72 s20,66 s192,38 10,66 s11,62 s383,19 5,66 s6,61 s772,93
32 40,71 s41,66 s192,36 21,04 s21,99 s382,43 10,97 s11,92 s763,17

Qwen 3.5 122B

Q4_K_M · approx. 78,7 GB VRAM per model instance

With 32 concurrently active queries, two GPUs achieved 1,98 times and four GPUs 4,00 times the total throughput of a single GPU.

Time to response start during parallel use

P95 time to the first visible response token · 1 to 32 concurrently active queries

The start time indicates the first visible response token.

Total throughput at 32 active queries

1 GPU
1,00×
2 GPUs
1,98×
4 GPUs
4,00×

Interpretation

With four concurrently active queries, the P95 response start was 10,60 seconds with X1, 6,11 seconds with X2, and 1,70 seconds with X4.

In this measurement series, the additional GPU capacity was converted almost entirely into parallel throughput.

Qwen 3.5 122B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries 1 GPU 2 GPUs 4 GPUs
P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s
1 1,69 s3,80 s66,89 1,72 s3,81 s66,89 1,69 s3,73 s67,72
2 4,25 s6,37 s79,28 1,70 s3,83 s132,13 1,69 s3,77 s134,60
4 10,60 s12,72 s78,40 6,11 s8,25 s157,73 1,70 s3,82 s266,73
8 23,36 s25,48 s79,12 12,56 s14,70 s156,00 7,84 s9,99 s304,79
16 48,79 s50,90 s80,45 25,45 s27,59 s157,47 14,59 s16,74 s312,79
32 99,57 s101,69 s78,67 51,04 s53,18 s155,73 27,16 s29,32 s314,95

Llama 4 Scout

Q4_K_M · approx. 66,2 GB VRAM per model instance

With 32 concurrently active queries, two GPUs achieved 1,97 times and four GPUs 3,91 times the total throughput of a single GPU.

Time to response start during parallel use

P95 time to the first visible response token · 1 to 32 concurrently active queries

The start time indicates the first visible response token.

Total throughput at 32 active queries

1 GPU
1,00×
2 GPUs
1,97×
4 GPUs
3,91×

Interpretation

With four concurrently active queries, the P95 response start was 12,93 seconds with X1, 6,98 seconds with X2, and 1,61 seconds with X4.

Even under high parallel load, scaling with one, two, and four complete model instances remained nearly linear.

Llama 4 Scout: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries 1 GPU 2 GPUs 4 GPUs
P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s
1 1,61 s4,54 s56,15 1,61 s4,49 s56,70 1,58 s4,43 s57,53
2 4,97 s7,93 s64,14 1,61 s4,55 s110,66 1,61 s4,47 s113,68
4 12,93 s15,87 s63,20 6,98 s9,96 s127,17 1,61 s4,58 s224,62
8 28,69 s31,64 s63,63 15,29 s18,27 s126,40 9,15 s12,15 s251,69
16 60,31 s63,26 s64,00 31,04 s34,03 s124,98 17,30 s20,31 s250,50
32 123,56 s126,51 s63,90 62,93 s65,91 s125,92 33,41 s36,42 s249,97

Mistral Medium 3.5 128B

Q4_K_M · approx. 79,5 GB VRAM per model instance · mean of 3 runs

In the more robust three-run test, with 32 active queries, two GPUs achieved 1,98 times and four GPUs 3,83 times the total throughput.

Time to response start during parallel use

P95 time to the first visible response token · 1 to 32 concurrently active queries

The start time indicates the first visible response token.

Total throughput at 32 active queries

1 GPU
1,00×
2 GPUs
1,98×
4 GPUs
3,83×

Interpretation

With four concurrently active queries, the P95 response start was 62,11 seconds with X1, 25,81 seconds with X2, and 3,61 seconds with X4.

The published Mistral values are based on the mean of three nearly identical measurement runs.

Mistral Medium 3.5 128B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries 1 GPU 2 GPUs 4 GPUs
P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s
1 3,52 s20,09 s12,67 3,42 s19,38 s13,21 3,39 s18,99 s13,60
2 22,80 s39,34 s12,80 3,55 s20,32 s24,55 3,47 s19,52 s26,43
4 62,11 s78,70 s12,80 25,81 s42,56 s25,60 3,61 s20,55 s51,20
8 140,78 s157,37 s12,80 65,38 s82,12 s25,60 32,95 s49,86 s49,01
16 298,14 s314,73 s13,01 144,75 s161,51 s25,60 74,12 s91,06 s49,23
32 612,52 s629,11 s12,94 303,37 s320,11 s25,60 153,50 s170,41 s49,56

DeepSeek R1 70B

Q4_K_M · approx. 43,7 GB VRAM per model instance · Thinking enabled

With 32 concurrently active queries, two GPUs achieved 1,97 times and four GPUs 3,81 times the total throughput of a single GPU.

Time to first generated token during parallel use

P95 including preceding reasoning output · 1 to 32 concurrently active queries

For this reasoning model, the start time indicates the first generated token; this may initially be part of the thinking phase.

Total throughput at 32 active queries

1 GPU
1,00×
2 GPUs
1,97×
4 GPUs
3,81×

Interpretation

With four concurrently active queries, the P95 time to the first generated token was 37,17 seconds with X1, 17,45 seconds with X2, and 2,54 seconds with X4.

For DeepSeek, the start time includes the preceding reasoning output and is therefore not part of the cross-model mean without Thinking.

DeepSeek R1 70B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries 1 GPU 2 GPUs 4 GPUs
P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s
1 2,67 s12,08 s21,47 2,60 s11,77 s21,60 2,42 s11,32 s22,30
2 13,86 s23,29 s21,60 2,56 s12,15 s41,60 2,50 s11,69 s44,59
4 37,17 s46,69 s21,83 17,45 s27,06 s43,20 2,54 s12,25 s85,88
8 83,67 s93,20 s21,66 41,58 s51,19 s42,42 23,33 s33,05 s82,07
16 177,09 s186,60 s21,87 88,16 s97,77 s43,32 48,07 s57,78 s83,20
32 362,99 s372,48 s21,78 182,24 s191,84 s42,84 95,22 s104,92 s83,04

GPT-OSS 120B

MXFP4 · approx. 62,4 GB VRAM per model instance · Thinking level “low”

With 32 concurrently active queries, two GPUs achieved 1,98 times and four GPUs 3,97 times the total throughput of a single GPU.

Time to first generated token during parallel use

P95 including preceding reasoning output · 1 to 32 concurrently active queries

For this reasoning model, the start time indicates the first generated token; this may initially be part of the thinking phase.

Total throughput at 32 active queries

1 GPU
1,00×
2 GPUs
1,98×
4 GPUs
3,97×

Interpretation

With four concurrently active queries, the P95 time to the first generated token was 5,60 seconds with X1, 3,29 seconds with X2, and 0,94 seconds with X4.

For GPT-OSS, the start time includes the preceding reasoning output and is therefore not part of the cross-model mean without Thinking.

GPT-OSS 120B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries 1 GPU 2 GPUs 4 GPUs
P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s P95 StartP95 durationTokens/s
1 1,10 s2,52 s104,87 0,93 s2,33 s109,00 0,96 s2,34 s108,18
2 2,08 s3,52 s142,86 0,93 s2,35 s216,36 0,93 s2,32 s219,66
4 5,60 s7,03 s144,51 3,29 s4,74 s279,93 0,94 s2,37 s432,71
8 12,64 s14,07 s142,40 7,39 s8,85 s284,88 5,20 s6,66 s564,84
16 26,72 s28,16 s144,29 14,71 s16,17 s282,39 9,73 s11,19 s555,99
32 54,87 s56,30 s143,64 29,13 s30,60 s284,69 17,22 s18,68 s569,58
Show benchmark methodology

Scaling was measured on the same Brainsystems X4 test platform with four NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs. For the individual measurement series, one, two, or four GPU-dedicated Ollama instances were active in sequence. One complete model instance ran per GPU.

GPUNVIDIA RTX PRO 6000 Blackwell Max-Q
GPU memory96 GB GDDR7 ECC per GPU
CPUAMD EPYC 9254
Worker architectureone complete model instance per GPU
Active GPUs1 / 2 / 4
Concurrently active queries1 / 2 / 4 / 8 / 16 / 32
Duration per measurement point300 seconds
Context window8.192 Tokens
Outputstandardized to 256 generated tokens
Measurement pathdirect Ollama API
Runs1 or 3 per model; multiple runs are shown as the arithmetic mean
Mean without ThinkingQwen 3 30B, Qwen 3.5 122B, Llama 4 Scout, and Mistral Medium 3.5 128B
Reasoning modelsDeepSeek R1 70B with Thinking=true; GPT-OSS 120B with Thinking level low
Completed queries28.950
Failed queries0

The named models were used exclusively as technical reference workloads and are not part of the scope of delivery. The measured values apply to the specified test platform, model version, quantization, context length, runtime environment, and benchmark methodology. They do not establish any entitlement to identical results in different customer environments and make no claim to completeness.