Team Ai
Modelpublic

thelamapi/next-codex

sourceHugging Facemitupdated 7mo agoView on Hugging Face
1likes22downloads
Model Card

30bcoder

πŸ’» Next-Codex (L846MoE)

Code your future with our models.

![License: MIT](https://opensource.org/licenses/MIT) ![Architecture: MoE]() ![HuggingFace](https://huggingface.co/Lamapi/next-codex) ![Discord](https://discord.gg/XgH4EpyPD2)


πŸ“– Overview

Next-Codex is a high-performance, specialized Mixture-of-Experts (MoE) Large Language Model designed specifically for code generation, debugging, and software engineering tasks.

Unlike traditional dense models, Next-Codex utilizes a sparse architecture with 30 Billion total parameters, but only activates 3 Billion parameters per token. This unique design allows it to deliver the deep reasoning capabilities of a massive model while maintaining the ultra-low latency and inference cost of a lightweight 3B model. It is fine-tuned on a massive corpus of code across 20+ programming languages, making it the most efficient coding assistant in its class.


⚑ Highlights

  • β€”πŸ‡ΉπŸ‡· TΓΌrkiye’s First Specialized MoE Coding Model: Designed for speed and precision.
  • β€”πŸš€ Hyper-Efficient Inference: Runs with 3B active parameters, enabling deployment on consumer GPUs (e.g., RTX 3090/4090).
  • β€”πŸ’» SOTA Coding Performance: Surpasses Claude Sonnet 4 and rivals o3-High in Python & JavaScript benchmarks.
  • β€”πŸŒ Polyglot Programming: Master-level proficiency in Python, JS/TS, Rust, Go, C++, SQL, and Swift.
  • β€”πŸ§  Context-Aware Debugging: Excellent at understanding large codebases and suggesting architectural improvements.
  • β€”πŸ’ Production Ready: Optimized for autocomplete, unit test generation, and docstring creation.

πŸ“Š Benchmark Performance (Coding & Logic)

Next-Codex achieves state-of-the-art results among open-weights coding models, balancing extreme efficiency with high accuracy.

Benchmarks are being conducted... ---

πŸš€ Installation & Usage

Note: Due to the MoE architecture, this model is memory efficient. You can run it comfortably on 24GB VRAM GPUs (4-bit quantization highly recommended for lower VRAM).

!pip install unsloth transformers
python
from unsloth import FastLanguageModel

# Load the MoE Model
model, tokenizer = FastLanguageModel.from_pretrained(
    "Lamapi/next-codex",
    load_in_4bit = True, # Optimized for 24GB VRAM
)

messages = [
    {"role": "system", "content": "You are Next-Codex, an expert software engineer and AI coding assistant."},
    {"role" : "user", "content" : "Write a highly optimized Rust function to calculate the Fibonacci sequence using memoization."}
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize = False,
    add_generation_prompt = True
)

from transformers import TextStreamer
_ = model.generate(
    **tokenizer(text, return_tensors = "pt").to("cuda"),
    max_new_tokens = 2048,
    temperature = 0.2, # Lower temperature for code precision
    top_p = 0.95,
    streamer = TextStreamer(tokenizer, skip_prompt = True),
)

🧩 Key Features

FeatureDescription
πŸ”€ Smart Routing (MoE)Dynamically routes tokens to the best "expert" layers, activating only 3B params for speed.
πŸ› οΈ Full-Stack MasteryTrained on frontend (React, Vue), backend (Django, Spring), and systems (C, Rust) code.
πŸ‡ΉπŸ‡· Code SupportExceptional ability to understand Turkish variable names and comments in legacy codebases.
🐞 Deep DebuggingAnalyzes stack traces and logic errors to provide instant fixes.
πŸ“ Docstring & TestingAutomatically generates Javadoc, PyDoc, and Unit Tests (Pytest/Jest).
πŸ”’ Secure CodingAligned to avoid common vulnerabilities (SQLi, XSS) in generated code.

πŸ“ Model Specifications

SpecificationDetails
ArchitectureMixture of Experts (MoE) Transformer
Total Parameters30 Billion
Active Parameters3 Billion (per token)
Context Window32k Tokens
Experts8 Experts (Top-2 Routing)
Training Data1T+ Tokens of Code (The Stack v2, GitHub, Synthetic)
QuantizationGGUF, AWQ, GPTQ supported

🎯 Ideal Use Cases

  • β€”IDE Autocomplete Plugins β€” Low latency makes it perfect for "Copilot" style completions.
  • β€”Legacy Code Refactoring β€” Converting outdated code to modern standards (e.g., Java 8 to Java 21).
  • β€”SQL Generation β€” Text-to-SQL for complex data analytics.
  • β€”Turkish/English Development β€” Teams working in bilingual environments.
  • β€”Algorithm Optimization β€” Reducing time complexity of existing functions.

πŸ“„ License

Licensed under the MIT License β€” free for commercial and non-commercial use.


πŸ“ž Contact & Support


Next-Codex β€” Smart as a giant, fast as a lightweight. The future of coding is MoE.

![Follow on HuggingFace](https://huggingface.co/Lamapi)