IFM's open-weight 7B model for coding, reasoning, and chat. Heabsy serves an FP8 deployment on a dedicated RTX 5090 with speculative decoding in the United States, with a 131K context window and an 8K output limit. Zero data retention is not verified for this deployment.
Heabsy’s Qwen3.8 deployment processes prompts and completions in memory within the European Economic Area without durable storage, logging, or training use. K2-Horizon-7B is served in the United States and does not have verified zero data retention. Its GDPR Art. 28 Data Processing Addendum is available at https://heabsy.com/dpa.
Added Sep 9, 2026
Context Window
131.1K
Max Output
8.2K
Input Price (Auto)
$0.050/1M
Output Price (Auto)
$0.20/1M
Cache Read (Auto)
$0.040/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare K2-Horizon-7B with similar models from the same provider or model family.
Agnes 3.0 Flash
agnes-3.0-flashAgnes 3.0 Flash is a low-cost model for coding, tool use, and multi-turn agent tasks. It supports text and image input, optional thinking, and a 512K-token context window.
Gemma 4 12B Semancer
gemma-4-12b-it-semancerGemma 4 12B Semancer is an open-weight finetune by UnstableLlama trained on a custom occult-philosophy dataset. It is designed for in-depth philosophical discussion of topics such as truth, free will, consciousness, and meaning, with answers developed from first principles. It supports image understanding, tool calling, optional reasoning, and a 131,072-token context window.
Gemma 4 12B StationKeeper
gemma-4-12b-it-station-keeperGemma 4 12B StationKeeper is an open-weight roleplay finetune with image understanding, tool calling, optional reasoning, and a 131,072-token context window.
Ling 3.0 Flash VL
inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.
Qwen 3.8 27B Queen
qwen/qwen3.8-27b-queenQwen 3.8 27B Queen is an open-weight roleplay finetune with image understanding, tool calling, optional reasoning, and a 262,144-token context window.
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This is a rate-limited beta with limited capacity, intended for testing rather than production use. Assume prompts and responses are logged by the provider and may be used for model training or service improvement. Do not send sensitive or confidential data.