In-house AI for companies – local on your own hardware
With a Brainsystems® in-house AI system, you can accelerate a wide range of time-consuming work processes: summarize documents, search company knowledge, pre-screen contracts, analyze tenders, evaluate technical documentation, and prepare decisions.
Processing takes place on your own hardware within your infrastructure. Internal documents, approved data sources, queries, and responses remain within the company – without external cloud processing and without the additional security, data protection, and compliance issues that can arise when confidential content is processed by third-party AI services.
Practical examples of how local AI can support you:
Use technical documentation
Brainsystems in-house AI can make technical manuals, datasheets, maintenance documents and internal documentation locally usable. Employees can ask concrete questions and receive structured answers based on approved sources.
Task for AI
Search the approved manuals and maintenance documents for error code E37. Summarize checks, possible causes and relevant spare parts.
AI thinking process
Ready
Classify error code
Search manuals
Condense checks
Map sources and spare parts
AI answer
Technical answer with sources
The AI finds relevant passages in manuals and maintenance documentation, condenses the checks and adds notes on affected assemblies and spare parts.
CauseSignal loss at position sensor likely
Service manual · page 42
CheckCheck connector X4 and measure sensor cable
Maintenance instruction · section 3.2
Spare partPosition sensor assembly PS-18
Spare parts list · item 18
NoteSimilar cases occurred after cable breaks in the drag-chain area
Maintenance reports · 2025
Technical verification remains with qualified experts; the AI shortens search and preparation.
Typical documents
PDF
Word
CAD documentation
Pre-review contracts
Brainsystems in-house AI can pre-structure contracts locally: obligations, deadlines, liability, termination, deviations from offers and notable clauses are extracted. Legal assessment deliberately remains with qualified professionals.
Task for AI
Map services, cooperation duties, deadlines, remuneration, liability and termination to their sources and flag ambiguities.
AI thinking process
Ready
Organize document set
Extract scope and duties
Compare fees and dates
Flag ambiguities
AI answer
Clause and scope overview
Scope, fee offer and contract draft are combined into a structured overview. Unclear mappings and deviations remain visible.
Special serviceFire protection coordination in contract, absent from fee offer
Contract §3.4 · bid page 6
Schedule deviationDetailed design due two weeks earlier than project plan
Contract annex 2 · project plan M4
ExpensesFlat rate of 4% clearly agreed
Contract §6.3
LiabilityReference to insurance sum is ambiguous
Contract §11.2 · annex 5
The structured pre-review does not replace legal or fee-regulation assessment.
Typical documents
PDF
Word
Analyze tenders
Brainsystems in-house AI can analyze tender documents locally and prepare them for processing. Requirements, deadlines, mandatory criteria, evidence, risks and open questions are structured into a usable overview.
Task for AI
Analyze the tender documents. Create an overview of deadlines, mandatory criteria, required evidence, technical requirements, risks and open questions.
AI thinking process
Ready
Capture documents
Identify deadlines and mandatory criteria
Collect evidence and risks
Create processing list
AI answer
Tender prepared
The AI creates a structured processing list with deadlines, evidence, mandatory requirements and points that should be clarified before submission.
DeadlineSubmission by May 14, questions by Apr 30
Tender documents · section 1
Mandatory criterionReference evidence for comparable projects required
Eligibility criteria · page 7
Open questionInterface description in annex 3 is ambiguous
Annex 3 · section 2.4
Next stepAlign evidence list and responsibilities
AI summary
The answer serves as a structured working basis for tender processing.
Typical documents
PDF
Word
Excel
Identify risks and gaps
Brainsystems in-house AI can search documents for risks, missing information, contradictory statements and unclear responsibilities. The result is not an automated decision, but a prioritized review list for experts.
Task for AI
Flag conflicting figures, missing explanations, unusual variances and unsupported assumptions.
AI thinking process
Ready
Synchronize data
Reconcile metrics
Detect anomalies
Document evidence
AI answer
Variance and clarification list
Figures and explanations are reconciled across report, spreadsheets and prior periods. Findings receive sources and clarification needs.
Figure conflictRevenue report: EUR 8.42m · sheet: EUR 8.31m
Report p. 4 · export cell F18
Unusual varianceService costs +27% month over month
Cost centers 410–430
Missing explanationEUR 180k provision without commentary
Balance sheet item 3.7
SupportedMaterial cost increase supported by documented price change
Supplier notice · 3 Jun 2026
The thinking machine flags review needs. Accounting assessment and approval remain with responsible professionals.
Typical documents
PDF
Excel
CSV
Prepare decisions
Brainsystems in-house AI can turn existing information into clear decision briefs. Options, criteria, risks, assumptions and open questions are structured so people can make faster and better decisions.
Task for AI
Summarize options, mandatory requirements, budget, risks and decisions still required.
AI thinking process
Ready
Capture need and target
Map constraints
Structure options
Name open decisions
AI answer
Structured procurement basis
Need, constraints and options are combined into a traceable working basis without pre-empting a procurement decision.
Mandatory requirementPersonal data must be processed locally
Data protection concept · section 3
Budget ceilingMaximum EUR 480,000 in current budget
Budget approval · item 8120
Open scope boundaryOperational support after year 1 is undefined
Specification · chapter 7
Next approvalIT security concept required before publication
Project plan · milestone M3
The presentation supports internal preparation and does not replace procurement, budget or data-protection review.
Typical documents
PDF
Word
Excel
Compare documents
Brainsystems in-house AI can compare offers, contract versions, requirements, specifications and document revisions locally. The AI highlights deviations, contradictions and open points as a basis for expert review.
Task for AI
Compare the bids with the bill of quantities. Show missing items, quantity differences, alternatives and reservations.
AI thinking process
Ready
Normalize documents
Match items
Classify deviations
Prioritize review
AI answer
Deviation matrix for bid review
Items are matched despite differing terminology, and deviations are structured by type and relevance.
Missing itemsBid B: 3 items not included
BoQ 02.14, 04.08 and 07.03
Quantity differenceBid C: item 05.12 lists 180 instead of 240 m²
Bid C · page 38
Alternative productBid A names an equivalent product alternative
Item 03.07 · data sheet A-17
CompletenessBid A covers all BoQ items
Automated item matching
The matrix is a working basis. Technical equivalence and procurement-law assessment remain human expert decisions.
Typical documents
PDF
Excel
Search company knowledge
Brainsystems in-house AI can search and evaluate internal documents, knowledge areas and approved databases locally. Employees ask concrete questions and receive understandable answers based on company knowledge — without cloud processing or data outflow.
Task for AI
Which checks are specified for error code E37, and which spare parts were used in similar cases?
AI thinking process
Ready
Decompose question
Search approved sources
Assess evidence
Compose sourced answer
AI answer
Source-based technical answer
Manual instructions and internal experience are combined in one answer without obscuring their origin.
Check 1Inspect connector X17 with power isolated
Service manual · page 214
Check 2Measure signal voltage at sensor S4
Service manual · page 215
Common spare partSensor S4, part 44-2187
Cases 2025-041 and 2026-008
NoteOne case reports a cable break rather than sensor failure
Maintenance report 2025-063
Sources remain visible. Technical work is performed only in accordance with applicable safety and work instructions.
Typical documents
PDF
Word
CAD documentation
Summarize documents
Brainsystems in-house AI can summarize reports, minutes, contracts and technical documents locally. Long documents become compact overviews with key statements, deadlines, open issues and traceable sources — without cloud processing or data outflow.
Task for AI
Create an objective overview of proposed decisions, costs, deadlines, open questions and differing positions.
AI thinking process
Ready
Capture papers
Extract decisions
Structure costs and dates
Compare positions
AI answer
Meeting brief at a glance
The extensive documents are condensed into a role-neutral overview of decisions, costs and open questions.
Decision needFour separate decisions required
Paper · section 6
BudgetEUR 420,000 one-off, EUR 38,000 annually
Financial annex · pages 2–4
Open questionResponsibility for recurring costs unclear
Statements A and C
Shared positionUnderlying need is undisputed
All three statements
The thinking machine structures the documents but does not make a political or legal decision.
Typical documents
PDF
Word
Scan
Why in-house AI?
Local. Secure. Independent.
In-house AI: data sovereignty and digital autonomy
In-house AI creates a dedicated AI infrastructure directly within the company. Models, queries, documents, and company data are processed in an environment you control. This makes AI viable for long-term use – for confidential information, core business processes, and the development of your own digital expertise.
Company data stays internal
Documents, queries, and generated responses are processed within your own infrastructure. Companies decide where data is stored, which systems can access information, and which users are allowed to use which content.
Your own security and access concepts
Brainsystems® can be integrated into existing data protection, security, and compliance structures – with defined data sources, role-based permissions, separate knowledge areas, and internal approval processes.
AI performance without token dependency
The available computing power can be used across the company. Frequent queries, extensive documents, and growing usage are processed within the available system capacity – without usage-based cloud or token subscriptions.
Operation within your own network
Local AI operates directly within the company environment. Employees and connected systems access an internally provided platform. Availability, maintenance windows, and priorities remain under your own control.
Company-specific pipelines and workflows
In-house AI can be integrated into your own processes step by step: from document search and internal knowledge queries to database, ERP, CRM, or ticketing system connections. This creates your own AI workflows that fit existing processes rather than being dictated by an external platform.
Example illustration: Local AI system with knowledge base, internal documents, and know-how - without the cloud.
Brainsystems® Product Family
Choose the right local AI for your business
Brainsystems® systems are powerful
on-premise AI platforms for businesses – from compact local AI
to large-memory platforms for extensive knowledge bases and
demanding AI workloads.
Brainsystems® 64G
Typical UseTeam, office, department, development
64GB ECC Pro Memory
for local AI, RAG, knowledge search and document analysis
1TB Local Storage
High-Speed NVMe Storage
AI Workloads
Ideal for local assistants, document search,
knowledge queries and clearly defined AI tasks
Multi-GPU Ready
scalable to up to 4 GPU accelerators
for more parallel AI requests
All Brainsystems® systems are based on
a professional, modular server platform with dedicated
AI acceleration and extensive expansion reserves.
28-Core CPU
AMD EPYC Pro Chip for parallel data processing,
RAG, embeddings and AI services
High-Speed AI
Dedicated GPU acceleration with high
AI Memory Bandwidth
Dual 10GbE
Two integrated 10-Gigabit Ethernet interfaces
for enterprise networks and High-Speed Storage
Integrated Remote Management
Remote access for system monitoring, diagnostics and administration
independent of the operating system
Modular by Design
Tower system for office use or 19" rack –
with extensive expansion options for future requirements
StorageCube Ready
Direct 10GbE connection to optional
StorageCube® for large local
data and knowledge bases
Brainsystems® is a local AI hardware platform.
The customer is responsible for the selection, license review, installation, configuration, and use of individual AI models.
AI models mentioned are used solely as a technical reference for memory requirements, runtime behavior, and system sizing.
They are not part of the standard scope of delivery and do not constitute a recommendation, approval, or assurance for production use.
System benchmarks · concurrent AI queries
More GPUs – More users
More GPUs do not reduce the response time of individual AI queries. They enable more concurrent AI users.
The appropriate Brainsystems® configuration level depends not only on the number of employees, but on how many AI tasks should start smoothly at the same time. The benchmark shows how 1, 2, and 4 GPUs affect response start under parallel use.
User experience is determined not only by the raw compute performance of a single query. In practice, what matters most is how the system behaves when several people use the in-house AI at nearly the same time: summarizing documents, comparing content, querying internal information, or starting longer analyses.
Assuming a desired response start of approximately 30 seconds, the benchmark gives the following picture:
5 parallel AI queries with 1 GPU -> 29,5s (to P95 response start)
10 parallel AI queries with 2 GPUs -> 31,7s (to P95 response start)
20 parallel AI queries with 4 GPUs -> 35,0s (to P95 response start)
It is important to distinguish between concurrently active AI queries and AI users. Not every user submits a query at exactly the same moment. The actual number of people who can work with a system can therefore be significantly higher. It depends on the type of tasks, frequency of use, document length, model size, and typical peak loads within the company.
In this context, more GPUs do not simply mean “more speed” for a single query. The main effect is additional parallel capacity: multiple model instances can work simultaneously, queues become shorter, and visible response start remains more controllable as usage increases. This is precisely what matters for day-to-day user experience.
Response start with 1, 2, and 4 GPUs
The chart shows the P95 response start for multiple concurrently active AI queries. This makes it clear when additional GPUs stabilize the user experience and keep parallel use more responsive.
X1 · 1 GPUX2 · 2 GPUsX4 · 4 GPUsReference point
The values shown are benchmark and reference values from a defined test configuration. The reference points of 5, 10, and 20 concurrent AI queries were interpolated between adjacent benchmark levels. In practice, the model, context length, document size, runtime environment, and usage behavior can affect the results.
AI model sizes and VRAM
Response quality
Model size matters. Context is crucial.
The available GPU memory helps determine which AI models can be run locally. In simplified terms:
The larger a model is, the better it can handle complex questions, long-range relationships, and tasks with many conditions.
For many business tasks, however, a medium-sized model that can run with 32 GB of VRAM is already sufficient.
In practice, however, model size is only one part of response quality. The right context is often more important:
internal documents, manuals, policies, contracts, technical documentation, project data, or approved knowledge areas.
This company context is what turns a general-purpose language model into a useful assistant for specific internal tasks.
A Brainsystems® in-house AI system can be connected to internal data sources with specifically controlled read-only access.
This gives the AI access to relevant company information without confidential content being sent to external cloud services
for processing. The data sources remain controllable: read-only, isolated, role-based, and tailored to the respective application.
Larger models become particularly relevant when very complex analyses, many conditions, long documents, or especially large
contexts need to be processed. The choice of models remains flexible: multiple models can be used in parallel,
switched, and tested. The existing internal documents, data sources, and company context remain available.
Named AI models, model names, product names, and trademarks belong to their respective rights holders.
They are named solely for the technical classification of model size, memory requirements, and local executability
under specific test and configuration conditions. Their mention does not constitute scope of delivery, a recommendation,
certification, ranking, or quality comparison of individual AI models, manufacturers, or providers.
AI models are not part of the standard scope of delivery of Brainsystems®.
Availability, license terms, and usage rights for individual models must be reviewed separately before use
within the respective company.
System benchmarks · parallel AI usage
Multi-GPU scaling with concurrent AI queries
The tabs show only measured hardware, response-time, and scaling values within the respective reference model.
They do not constitute a quality comparison, ranking, or recommendation of individual AI models.
The results show the behavior of the local system platform under defined test conditions.
DeepSeek R1 and GPT-OSS are not included in the cross-model mean without Thinking because of their reasoning output.
Qwen 3 30B
Q4_K_M · approx. 19,2 GB VRAM per model instance
With 32 concurrently active queries, two GPUs achieved 1,99 times and four GPUs 3,97 times the total throughput of a single GPU.
Time to response start during parallel use
P95 time to the first visible response token · 1 to 32 concurrently active queries
The start time indicates the first visible response token.
Total throughput at 32 active queries
1 GPU
1,00×
2 GPUs
1,99×
4 GPUs
3,97×
Interpretation
With four concurrently active queries, the P95 response start was 4,28 seconds with X1, 2,49 seconds with X2, and 0,70 seconds with X4.
Under the highest measured load, X4 achieved almost four times the total throughput of a single GPU.
Qwen 3 30B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries
1 GPU
2 GPUs
4 GPUs
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
1
0,71 s
1,66 s
153,60
0,70 s
1,64 s
155,25
0,74 s
1,69 s
152,77
2
1,64 s
2,58 s
193,24
0,70 s
1,64 s
308,02
0,70 s
1,65 s
308,85
4
4,28 s
5,24 s
194,06
2,49 s
3,43 s
385,64
0,70 s
1,64 s
616,87
8
9,42 s
10,37 s
198,19
5,24 s
6,19 s
388,95
3,07 s
4,02 s
765,50
16
19,72 s
20,66 s
192,38
10,66 s
11,62 s
383,19
5,66 s
6,61 s
772,93
32
40,71 s
41,66 s
192,36
21,04 s
21,99 s
382,43
10,97 s
11,92 s
763,17
Qwen 3.5 122B
Q4_K_M · approx. 78,7 GB VRAM per model instance
With 32 concurrently active queries, two GPUs achieved 1,98 times and four GPUs 4,00 times the total throughput of a single GPU.
Time to response start during parallel use
P95 time to the first visible response token · 1 to 32 concurrently active queries
The start time indicates the first visible response token.
Total throughput at 32 active queries
1 GPU
1,00×
2 GPUs
1,98×
4 GPUs
4,00×
Interpretation
With four concurrently active queries, the P95 response start was 10,60 seconds with X1, 6,11 seconds with X2, and 1,70 seconds with X4.
In this measurement series, the additional GPU capacity was converted almost entirely into parallel throughput.
Qwen 3.5 122B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries
1 GPU
2 GPUs
4 GPUs
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
1
1,69 s
3,80 s
66,89
1,72 s
3,81 s
66,89
1,69 s
3,73 s
67,72
2
4,25 s
6,37 s
79,28
1,70 s
3,83 s
132,13
1,69 s
3,77 s
134,60
4
10,60 s
12,72 s
78,40
6,11 s
8,25 s
157,73
1,70 s
3,82 s
266,73
8
23,36 s
25,48 s
79,12
12,56 s
14,70 s
156,00
7,84 s
9,99 s
304,79
16
48,79 s
50,90 s
80,45
25,45 s
27,59 s
157,47
14,59 s
16,74 s
312,79
32
99,57 s
101,69 s
78,67
51,04 s
53,18 s
155,73
27,16 s
29,32 s
314,95
Llama 4 Scout
Q4_K_M · approx. 66,2 GB VRAM per model instance
With 32 concurrently active queries, two GPUs achieved 1,97 times and four GPUs 3,91 times the total throughput of a single GPU.
Time to response start during parallel use
P95 time to the first visible response token · 1 to 32 concurrently active queries
The start time indicates the first visible response token.
Total throughput at 32 active queries
1 GPU
1,00×
2 GPUs
1,97×
4 GPUs
3,91×
Interpretation
With four concurrently active queries, the P95 response start was 12,93 seconds with X1, 6,98 seconds with X2, and 1,61 seconds with X4.
Even under high parallel load, scaling with one, two, and four complete model instances remained nearly linear.
Llama 4 Scout: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries
1 GPU
2 GPUs
4 GPUs
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
1
1,61 s
4,54 s
56,15
1,61 s
4,49 s
56,70
1,58 s
4,43 s
57,53
2
4,97 s
7,93 s
64,14
1,61 s
4,55 s
110,66
1,61 s
4,47 s
113,68
4
12,93 s
15,87 s
63,20
6,98 s
9,96 s
127,17
1,61 s
4,58 s
224,62
8
28,69 s
31,64 s
63,63
15,29 s
18,27 s
126,40
9,15 s
12,15 s
251,69
16
60,31 s
63,26 s
64,00
31,04 s
34,03 s
124,98
17,30 s
20,31 s
250,50
32
123,56 s
126,51 s
63,90
62,93 s
65,91 s
125,92
33,41 s
36,42 s
249,97
Mistral Medium 3.5 128B
Q4_K_M · approx. 79,5 GB VRAM per model instance · mean of 3 runs
In the more robust three-run test, with 32 active queries, two GPUs achieved 1,98 times and four GPUs 3,83 times the total throughput.
Time to response start during parallel use
P95 time to the first visible response token · 1 to 32 concurrently active queries
The start time indicates the first visible response token.
Total throughput at 32 active queries
1 GPU
1,00×
2 GPUs
1,98×
4 GPUs
3,83×
Interpretation
With four concurrently active queries, the P95 response start was 62,11 seconds with X1, 25,81 seconds with X2, and 3,61 seconds with X4.
The published Mistral values are based on the mean of three nearly identical measurement runs.
Mistral Medium 3.5 128B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries
1 GPU
2 GPUs
4 GPUs
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
1
3,52 s
20,09 s
12,67
3,42 s
19,38 s
13,21
3,39 s
18,99 s
13,60
2
22,80 s
39,34 s
12,80
3,55 s
20,32 s
24,55
3,47 s
19,52 s
26,43
4
62,11 s
78,70 s
12,80
25,81 s
42,56 s
25,60
3,61 s
20,55 s
51,20
8
140,78 s
157,37 s
12,80
65,38 s
82,12 s
25,60
32,95 s
49,86 s
49,01
16
298,14 s
314,73 s
13,01
144,75 s
161,51 s
25,60
74,12 s
91,06 s
49,23
32
612,52 s
629,11 s
12,94
303,37 s
320,11 s
25,60
153,50 s
170,41 s
49,56
DeepSeek R1 70B
Q4_K_M · approx. 43,7 GB VRAM per model instance · Thinking enabled
With 32 concurrently active queries, two GPUs achieved 1,97 times and four GPUs 3,81 times the total throughput of a single GPU.
Time to first generated token during parallel use
P95 including preceding reasoning output · 1 to 32 concurrently active queries
For this reasoning model, the start time indicates the first generated token; this may initially be part of the thinking phase.
Total throughput at 32 active queries
1 GPU
1,00×
2 GPUs
1,97×
4 GPUs
3,81×
Interpretation
With four concurrently active queries, the P95 time to the first generated token was 37,17 seconds with X1, 17,45 seconds with X2, and 2,54 seconds with X4.
For DeepSeek, the start time includes the preceding reasoning output and is therefore not part of the cross-model mean without Thinking.
DeepSeek R1 70B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Concurrently active queries
1 GPU
2 GPUs
4 GPUs
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
P95 Start
P95 duration
Tokens/s
1
2,67 s
12,08 s
21,47
2,60 s
11,77 s
21,60
2,42 s
11,32 s
22,30
2
13,86 s
23,29 s
21,60
2,56 s
12,15 s
41,60
2,50 s
11,69 s
44,59
4
37,17 s
46,69 s
21,83
17,45 s
27,06 s
43,20
2,54 s
12,25 s
85,88
8
83,67 s
93,20 s
21,66
41,58 s
51,19 s
42,42
23,33 s
33,05 s
82,07
16
177,09 s
186,60 s
21,87
88,16 s
97,77 s
43,32
48,07 s
57,78 s
83,20
32
362,99 s
372,48 s
21,78
182,24 s
191,84 s
42,84
95,22 s
104,92 s
83,04
GPT-OSS 120B
MXFP4 · approx. 62,4 GB VRAM per model instance · Thinking level “low”
With 32 concurrently active queries, two GPUs achieved 1,98 times and four GPUs 3,97 times the total throughput of a single GPU.
Time to first generated token during parallel use
P95 including preceding reasoning output · 1 to 32 concurrently active queries
For this reasoning model, the start time indicates the first generated token; this may initially be part of the thinking phase.
Total throughput at 32 active queries
1 GPU
1,00×
2 GPUs
1,98×
4 GPUs
3,97×
Interpretation
With four concurrently active queries, the P95 time to the first generated token was 5,60 seconds with X1, 3,29 seconds with X2, and 0,94 seconds with X4.
For GPT-OSS, the start time includes the preceding reasoning output and is therefore not part of the cross-model mean without Thinking.
GPT-OSS 120B: measured P95 start time, standardized duration for 256 generated tokens, and total throughput.
Scaling was measured on the same Brainsystems X4 test platform with four NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs. For the individual measurement series, one, two, or four GPU-dedicated Ollama instances were active in sequence. One complete model instance ran per GPU.
GPU
NVIDIA RTX PRO 6000 Blackwell Max-Q
GPU memory
96 GB GDDR7 ECC per GPU
CPU
AMD EPYC 9254
Worker architecture
one complete model instance per GPU
Active GPUs
1 / 2 / 4
Concurrently active queries
1 / 2 / 4 / 8 / 16 / 32
Duration per measurement point
300 seconds
Context window
8.192 Tokens
Output
standardized to 256 generated tokens
Measurement path
direct Ollama API
Runs
1 or 3 per model; multiple runs are shown as the arithmetic mean
Mean without Thinking
Qwen 3 30B, Qwen 3.5 122B, Llama 4 Scout, and Mistral Medium 3.5 128B
Reasoning models
DeepSeek R1 70B with Thinking=true; GPT-OSS 120B with Thinking level low
Completed queries
28.950
Failed queries
0
The named models were used exclusively as technical reference workloads and are not part of the scope of delivery.
The measured values apply to the specified test platform, model version, quantization, context length, runtime environment, and benchmark methodology.
They do not establish any entitlement to identical results in different customer environments and make no claim to completeness.