Skip to content

Add Mixat dataset (Emirati-English code-switched speech) - #713

Merged
zaidalyafeai merged 3 commits into
ARBML:mainfrom
kftadros:add-mixat-dataset
Jul 27, 2026
Merged

Add Mixat dataset (Emirati-English code-switched speech)#713
zaidalyafeai merged 3 commits into
ARBML:mainfrom
kftadros:add-mixat-dataset

Conversation

@kftadros

Copy link
Copy Markdown
Contributor

Adds the Mixat corpus: 15 hours of Emirati Arabic–English code-switched speech collected from two public podcasts featuring native Emirati speakers, with manual transcriptions (Al Ali & Aldarmaki, LREC-COLING 2024 / SIGUL).

Validated locally against schema.json with validate_schema.py (passed). Recorded the venue as LREC-COLING since SIGUL isn't in the controlled venue list; happy to adjust any fields on request.

@zaidalyafeai

Copy link
Copy Markdown
Contributor

lgtm @kftadros , thank you.

@zaidalyafeai
zaidalyafeai merged commit 43d4506 into ARBML:main Jul 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants