Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

TII launches Falcon-Emirati-7B for Emirati Arabic

Technology Innovation Institute launched Falcon-Emirati-7B, an adaptation of Falcon-H1-Arabic for Emirati speech and cultural context, without disclosing key training-data details.

D
Oct 6, 2026 · 2 min read

Technology Innovation Institute launched Falcon-Emirati-7B, a 7-billion-parameter language model adapted to understand and generate Emirati Arabic. TII says it is available through Falcon Chat, but the announcement did not identify a downloadable repository or license.

According to TII, the model was adapted from the 7B version of Falcon-H1-Arabic rather than trained from scratch. The base family combines Mamba state-space components with Transformer attention in parallel inside each block. In practical terms, the design pairs a mechanism intended to process long sequences efficiently with attention that can directly connect relevant parts of the input. TII lists a 256,000-token context window for the 7B base model.

TII frames the main adaptation challenge as a shortage of written Emirati Arabic because the dialect is primarily spoken. Its pipeline combined curated dialect text from Emirati websites and forums, Modern Standard Arabic material about Emirati culture, and synthetic Emirati text constrained by dialect glossaries, dictionaries and style rules.

The institute says it varied the amount and training stage of dialect data, the balance of authentic and synthetic text, and the amount of MSA cultural material. Automatic scoring was paired with reviews by native Emirati speakers. TII did not disclose corpus sizes, token counts, collection dates, source domains, filtering rules or train-validation contamination checks. Nor did it provide enough detail to reproduce the continued-pretraining, supervised-fine-tuning or preference-optimization schedule.

For evaluation, TII used Alyah, its benchmark of 1,173 four-option questions across seven categories. TII says native Emirati speakers collected the questions manually, while language models generated distractor answers that were reviewed. The Alyah dataset is available separately.

TII reports 84.83% accuracy for Falcon-Emirati-7B on Alyah, leading the Arabic and multilingual models in its selected comparison. That comparison excluded Falcon-H1-Arabic models because Falcon-Emirati was built from the family. TII produced the result, which has not been independently reproduced.

The institute also tested open-ended responses to the same 1,173 questions, using Gemini 3.7 Flash to judge correctness and Emirati-dialect fidelity. TII reports a partial-credit dialect-fidelity score of 0.52 for Falcon-Emirati-7B, compared with 0.05 or less for each of four comparison models. The release did not provide the full judge prompt, sampling settings, repeated-run variance, human-agreement measurements or uncertainty intervals.

In a separate cultural test, TII used 283 UAE scenarios from ArabCulture-Dialogue, an independently published dataset spanning 13 Arabic-speaking countries, 12 daily-life topics and 54 subtopics in MSA and national dialects. The benchmark’s authors found that tested models performed worse in dialectal settings than in MSA across three tasks. TII reports 85.57% multiple-choice accuracy for Falcon-Emirati-7B on the UAE subset, ahead of three models it tested. That score is TII’s result, not a finding from the benchmark paper.

TII cautions that Falcon-Emirati-7B may reproduce biases in its training data and make mistakes on rare expressions, localized references and other data-sparse cases. It recommends use-case-specific evaluation before sensitive or high-stakes use.

More news