Research · work in progress

AHMAD-1

An Arabic-first reasoning model, being built in the open by Ahmad A M Odeh. This page reports what has been measured, including the results that came back negative. No model has been released and the model card on file is a draft.

What it is

The goal is a small model that reasons in Arabic and understands the language’s root-and-pattern structure, trained with reinforcement learning against checkable rewards (for example, “is this the correct trilateral root?”). An optional geometric layer represents each concept as a subspace, not a point. The design rule is that no geometric component earns its place without a pre-registered kill test.

Measured so far (2026-09-19)

Base-model bake-off — 85 Arabic prompts (50 root extraction, 20 reasoning, 15 morphology), greedy decoding

ModelRootReasoningMorphologyOverall
Qwen3-0.6B-Base32.0%85.0%26.7%43.5%
Qwen3-4B-Base62.0%100%33.3%65.9%

Root accuracy doubles with scale; verb-form morphology stays flat (27–33%), which points to a training-signal gap rather than a capacity gap — a good target for reinforcement learning. A first run scored the 4B model’s reasoning at 5%; that turned out to be a scoring bug (the parser read an invented follow-up question), was fixed, and the corrected numbers above are the ones that count. The Falcon-H1-Arabic-7B candidate has not been scored: its repositories are gated.

Geometry probe — is a concept a point or a subspace? (Gemma-4-E2B-it, 90 word-in-context vectors)

  • Point / simplex thesis: rejected. Concept centroids sit at cosine ≈ 0.93, nowhere near an equiangular simplex (−0.5).
  • Subspace structure: real. A concept needs 2–3 dimensions (top direction explains about 54% of variance); principal angles between concepts average 21° against 88° for random — clearly separated.
  • Cross-lingual alignment: null. Arabic subspaces align with their English equivalents no better (48.9°) than with mismatched concepts (47.3°). Only transliterated Arabic lines up, so script — not meaning — drives the gap.

Verdict: conditional go. A monolingual subspace layer has measured structure behind it; the cross-lingual motivation is null and is not claimed. This shows the substrate exists, not that the layer will help.

What is blocked

The 20 papers it stands on

#PaperWhy it matters here
1The Geometry of Algorithms with Orthogonality Constraints
Edelman, Arias, Smith, 1998, SIAM J. Matrix Anal. Appl.
The foundational reference for Stiefel/Grassmannian geometry: geodesics, principal angles, exponential/log maps — the math layer under the new concept layer.
2Optimization Algorithms on Matrix Manifolds
Absil, Mahony, Sepulchre, 2008, Princeton UP
Riemannian SGD/CG/trust-region on G(k, R^d): the optimizer family for training subspace-constrained concept parameters without leaving the manifold.
3Efficient Algorithms for Inferences on Grassmann Manifolds
Gallivan, Srivastava, Liu, Van Dooren, 2003, IEEE SSP
O(n·k²) algorithms for geodesics and principal angles — makes per-token Grassmannian distance computation during inference cheap enough for real-time anomaly scoring.
4Grassmann Discriminant Analysis: A Unifying View on Subspace-Based Learning
Hamm & Lee, 2008, ICML
Classification by projection/Binet–Cauchy kernels over principal angles — the template for mapping a latent state to the nearest root-subspace ("which concept is active").
5Building Deep Networks on Grassmann Manifolds
Huang et al., 2018, AAAI
GrNet shows end-to-end layers can operate natively on Grassmannian points with mapping/projection/pooling blocks — an architectural pattern for the geometric concept layer.
6Riemannian Approach to Batch Normalization
Cho & Lee, 2017, NeurIPS
Demonstrates stochastic Grassmannian SGD with momentum inside ordinary deep training — proof manifold-constrained parameters can be optimized stably alongside Euclidean ones.
7Infeasible Deterministic, Stochastic, and Variance-Reduction Algorithms for Optimization under Orthogonality Constraints
Ablin & Peyré, 2023 · arXiv:2303.16510
The "landing method" makes orthogonality-constrained SGD practical at network scale without expensive retractions — the right optimizer if concept bases are trained jointly with a transformer.
8The Role of Principal Angles in Subspace Classification
Huang, Qiu, Calderbank, 2016, IEEE Trans. Signal Processing
Proves misclassification probability between subspaces is governed (low noise: product of sines; high noise: sum of squares of sines of principal angles) — the theoretical basis for using principal angles as the ISNAD-grade distance metric.
9Subspace Tracking with Dynamical Models on the Grassmannian
Saad-Falcon et al., 2023 (IEEE TSP / arXiv)
Tracks time-varying subspaces under distribution shift with geodesic/chordal regularizers — directly the "training-distribution subspace drifts; flag queries outside the neighborhood" mechanism of the 8200 framework.
10ViM: Out-of-Distribution with Virtual-Logit Matching
Wang et al., 2022, CVPR
OOD score = residual after projecting features onto the principal ID subspace of the classifier — a one-file implementation pattern for AHMAD-1's anomaly head.
11Out-of-Distribution Detection with Deep Nearest Neighbors
Sun, Ming, Zhu, Li, 2022, ICML
Non-parametric kNN on penultimate features beats parametric Mahalanobis with no distributional assumptions — a zero-training baseline and fallback for the anomaly product.
12Activation Subspaces for Out-of-Distribution Detection (ActSub)
Zöngür et al., 2025, ICCV
State of the art among subspace-based OOD scores using decisive/insignificant activation subspaces — the current quality bar the Grassmannian detector must beat.
13DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao et al., 2024 · arXiv:2402.03300
Introduces GRPO, the critic-free RL algorithm that halves PPO memory — the core training engine specified in the PRD.
14DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI (Guo et al.), 2025 · arXiv:2501.12948
Establishes the base→RL-with-verifiable-rewards recipe and its distills (1.5B–70B) — the strategic template for AHMAD-1's pipeline and the bar for "reasoning model" claims.
15SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Zeng et al., 2025 · arXiv:2503.18892
Systematic small-scale GRPO study: simple +1/0 correctness rewards, difficulty-matched data, and the warning that format rewards can *hurt* base-model exploration — the hyperparameter playbook for our 0.6B–4B runs.
16DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL
Luo et al., 2025 (tech report)
Iterative context-length scaling (8K→24K) with GRPO took a 1.5B model to 43.1% AIME — the reference point for what "domain expert at small scale" actually costs in compute.
17Qwen3 Technical Report
Yang et al. (Qwen Team), 2025 · arXiv:2505.09388
Documents Qwen3-0.6B/1.7B/4B with thinking/non-thinking modes (Apache-2.0) — our named base models (d=1024 for 0.6B, d=2560 for 4B) and their pre-RL baselines.
18SmolLM3: smol, multilingual, long-context reasoner
Allal et al. (Hugging Face), 2025
Fully-open 3B Apache-2.0 model trained on 11.2T tokens with dual think/no-think modes — the strongest fully-reproducible sub-4B training recipe in the literature and the methodology benchmark for "publishable in 30 days."
19The Era of 1-bit LLMs: All Large Language Models Are in 1.58 Bits
Ma et al., 2024 · arXiv:2402.17764
BitNet b1.58 ternary training matches FP16 at ≥3B with 3.55× less memory — the only credible path to genuinely sub-1GB *with* reasoning intact, but requires training-from-scratch (QAT), not post-hoc quantization.
20Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization on Llama-3.1-8B-Instruct
O'Neal et al., 2026 · arXiv:2601.14277
Controlled data on GGUF tiers: sub-3-bit (Q3_K_S) collapses instruction-following benchmarks; 4-bit K-quants stay near-FP16 — sets the Q2_K "compression floor" risk line for §5.4.

Citations were checked against live sources when the review was written (2026-09-19); anything unverifiable is marked as such in the source review.

Related: Arabic and open benchmarks, live · what is ready · AI Pulse