Führende Open-Source KI-Modelle sofort einsatzbereit

Wähle aus den leistungsstärksten State-of-the-Art Modellen. Vollständig optimiert für verteiltes Inferenz-Streaming mit maximalem Token-Durchsatz und 100% OpenAI Drop-In API.

💻 Coding & Agent Flagship 128k Context

Qwen 2.5 Coder 32B

Top-Tier Code-Generierung, Refactoring, Multi-File-Synthese und autonome Agenten-Steuerung.

Parameter32.7B
Min. VRAM20 GB
Speed~68 tok/s
Preis$0.25 / 1M
🧠 Reasoning & Math 64k Context

DeepSeek-R1 (Distill 32B / 671B MoE)

Reinforcement-Learning Reasoning-Modell mit explizitem Chain-of-Thought für komplexe Logik- und Mathe-Aufgaben.

Parameter32.0B / 671B
Min. VRAM22 GB
Speed~45 tok/s
Preis$0.35 / 1M
🏢 Enterprise LLM 128k Context

Llama 3.3 70B Instruct

Meta's Flaggschiff für mehrsprachige Dialoge, komplexe Analysen und professionelle Enterprise-Assistenten.

Parameter70.6B
Min. VRAM42 GB
Speed~52 tok/s
Preis$0.45 / 1M
👁️ Vision & Multimodal 32k Context

Qwen 2.5 VL 7B / MiniCPM-V

Visuelle Dokumentenanalyse, OCR, UI-Screenshots, Diagramme und detaillierte Bildbeschreibungen.

Parameter7.6B
Min. VRAM8 GB
Speed~85 tok/s
Preis$0.18 / 1M
⚡ Ultra High-Speed 128k Context

Mistral NeMo 12B

Gemeinsam mit NVIDIA trainiert. Extrem effiziente und reaktionsschnelle Allround-Inferenz.

Parameter12.2B
Min. VRAM10 GB
Speed~94 tok/s
Preis$0.15 / 1M
📱 Lightweight & Edge 32k Context

Phi-3 Mini (3.8B)

Microsofts ultra-kompaktes Modell. Perfekt für Edge-Nodes, Smartphones und schnelle Chat-Routinen.

Parameter3.8B
Min. VRAM4 GB
Speed~140 tok/s
Preis$0.09 / 1M

Jedes Modell mit 2 Zeilen Code ansteuern

Tausche einfach die `base_url` und den `api_key` in deinen bestehenden OpenAI SDKs aus:

inference_example.py
from openai import OpenAI

client = OpenAI(
    base_url="https://mesh.inetconnector.com/v1",
    api_key="cm_live_your_key"
)

response = client.chat.completions.create(
    model="qwen2.5:7b",
    messages=[{"role": "user", "content": "Schreibe einen Algorithmus für verteiltes Inferenz-Routing."}],
    stream=True
)

for chunk in response:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)