ZimanKitGitHub

Small toolkit.
Clear boundaries.

What ZimanKit changes, what it preserves, and how to make it part of your own project. Free to use, inspect, and build on.

Get started

  1. Choose Sorani, Kurmanji, or Unicode only. The script detector does not choose a language for you.
  2. Paste text or import a UTF-8 .txt file. Choose Normalize, Clean, or experimental Transliterate.
  3. Review the result, change report, and processing notes. Copy the result, download text, or export the full JSON report.

The playground runs entirely in your browser. Text is not submitted to the API, stored in cookies, or sent to analytics. The optional public API is a separate way to use the same engine.

Open the playground

Explicit language profiles

All operations first apply Unicode NFC. It composes canonically equivalent sequences, such as a letter plus a combining accent. ZimanKit deliberately avoids global NFKC, case folding, diacritic removal, and spelling correction.

ProfileNormalizationPreserved
Soraniك → ک and ي → ی, when the variant option is enabled.Distinct Kurdish letters, vowels, diacritics, digits, and original spelling.
KurmanjiNFC, including decomposed ç, ê, î, ş, and û.Capitalization and diacritics. Turkish ı is not silently changed to i.
Unicode onlyNFC only; no Kurdish character replacement.Arabic quotations and mixed-language spellings.

Sorani replacements apply to every matching character in the selected passage, including Arabic quotations. Use Unicode only for those passages. Similar-looking characters such as ه / ە, ر / ڕ, and ل / ڵ remain distinct.

Legacy font encodings and Arabic presentation forms are not converted in v0.1. Presentation forms are flagged. Script detection reports Arabic, Latin, mixed, or unknown based on letters; it does not establish whether text is Kurdish.

Cleaning, with choices

Clean applies NFC and the selected formatting options. It does not also apply Sorani variant replacements; use Normalize first if you want both.

OptionDefaultBehavior
Repeated spacesOnCollapse ordinary spaces and tabs. Preserve newlines and non-breaking spaces.
Line endingsOnConvert CRLF and CR to LF. Keep paragraphs and blank lines.
Trim line edgesOffRemove ordinary spaces and tabs at the beginning and end of each line.
Punctuation spacingOffRemove ordinary spaces before , ، ; ؛ ! ? ؟. Does not infer missing spaces or rewrite punctuation.
Elongation marksOffRemove tatweel (ـ).
Formatting controlsOffRemove U+200B, U+200E–200F, U+202A–202E, U+2060, U+2066–2069, and U+FEFF. Review direction and word boundaries afterwards.

Joiners (ZWJ), non-joiners (ZWNJ), combining marks, emoji sequences, and non-breaking spaces are preserved. Cleaning code or formatted documents may alter indentation; disable repeated spaces and trim options for those inputs.

Experimental transliteration

Transliteration changes the script, not the language. Sorani does not become Kurmanji. This feature produces a draft for review, not guaranteed pronunciation or translation.

ZimanKit uses a documented extended Hawar-style Latin mapping: ç, ş, j, x, ê, î, û, plus ł for ڵ, ř for ڕ, ḥ for ح, ẍ for غ, and ʿ for ع. This is an explicit project convention, not a claim of a universal Kurdish romanization standard.

Arabic → Latin: an annotated draft

The letters و and ی may represent multiple sounds. Single occurrences become {w/u} and {y/î}, rather than a guessed reading. Double وو becomes û. Initial ئ is removed only before a recognized written vowel. Unwritten short i (bizroke) is not reconstructed.

کوردی → k{w/u}rd{y/î}
ئاسۆ → aso
ڕێگا → řêga

Latin → Arabic: a lossy draft

Uses Sorani-style Arabic spelling for either selected variety. Initial written vowels receive a carrier; capitalization is lost, short i is omitted, w/u merge as و, and y/î merge as ی. Round trips are not guaranteed. Names, foreign words, URLs, and dialect-specific conventions need manual review. Unknown characters are retained and flagged.

No dictionary, AI model, or paid service is used. There is no independently measured linguistic accuracy score for this release.

Counts with defined meanings

  • Words: runs of Unicode letters or numbers, including combining marks. Internal apostrophes and joining controls remain part of a word. Hyphens and punctuation split words; decimals contribute two numeric runs.
  • Characters: grapheme clusters via Intl.Segmenter, including whitespace. A joined emoji can be one character but several code points. Unicode segmentation depends on the runtime’s Unicode version.
  • Sentences: approximate segments containing letters/numbers separated by . ! ? ؟ ۔ or …. Decimal points are excluded. Abbreviations, URLs, and unusual punctuation may inflate counts.
  • Changes: the sum of changed spans across pipeline rules. This is not an edit distance. The JSON report includes up to four before/after examples per rule.

Counts describe the result. Empty input has zero words, characters, sentences, and lines. Source and output remain available in the JSON report.

One engine. Your project.

The TypeScript library has no framework or service dependency. Build it from the repository, pack it, and install the local archive in your own app. v0.1 is available from GitHub; it has not been published to npm.

git clone https://github.com/zanaamed/zimankit.git
cd zimankit
npm ci
npm run build:core
npm pack --workspace zimankit
# In your other project:
npm install /path/to/zimankit/zimankit-0.1.0.tgz
import { processText, detectScript, getStatistics } from 'zimankit';

const result = processText('كوردی', {
  language: 'sorani',
  operation: 'normalize',
});
console.log(result.text); // کوردی
console.log(result.changes);
console.log(detectScript('Silav!').kind); // latin

processText(text, options) defaults to Unicode-only normalization. It returns the original, processed text, profile, script information, statistics, changes, warnings, experimental flag, and engine version. Invalid options throw TypeError; oversized text throws RangeError. Input must be well-formed Unicode and at most 20,000 UTF-16 code units.

A small REST API

Same-origin endpoint: POST /api/v1/process. Set Content-Type: application/json. The API defaults to Unicode-only normalization, supports CORS, and does not store request text in application code. Your text is sent to the server when you call this endpoint; the playground never calls it.

fetch('/api/v1/process', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    text: 'كوردی',
    options: { language: 'sorani', operation: 'normalize' },
  }),
}).then(response => response.json());

Supported options: language, operation, direction, kurdishVariants, collapseSpaces, normalizeLineEndings, trimLines, punctuationSpacing, removeTatweel, and removeControls. Flags are booleans. Transliterating requires a language and explicit arabic-to-latin or latin-to-arabic direction.

API limits: 2,000 UTF-16 code units per text and 100,000 bytes per JSON body. The smaller API limit helps bound free-plan CPU use; the playground and local library accept 20,000. Responses: 200 success, 400 invalid JSON/options, 405 wrong method, 413 oversized input, 415 wrong media type. Errors use { "error": "message" }. GET /api/v1/health returns status, version, and text limit. API availability is subject to the host’s free-plan quotas.

For local API testing, build the project and run npm run preview. The Next.js development server serves the playground; the API runs in the Cloudflare Worker preview.

Download it. Build it yourself.

Use Node.js 24 (or a supported Node version ≥22) and npm. No API key, account, database, or cloud subscription is needed to run the library and playground locally.

git clone https://github.com/zanaamed/zimankit.git
cd zimankit
npm ci
npm run dev
# Playground: http://127.0.0.1:3000

npm run check
npm run preview
# Playground and API: http://localhost:8787

Next.js exports static files into out/. These can be hosted on any static host; the API requires the included Worker. To deploy both to your own Cloudflare Workers Free account, run npx wrangler login, then npm run deploy. No storage or paid bindings are configured.

Improve the rules, together

This release has engineering tests, but still needs wider native-speaker review. To report a linguistic issue, include a short reproducible input, expected output, language/profile, and the reason or orthographic reference. Do not submit confidential text.

Source code is MIT licensed. Bundled Geist and Noto Sans Arabic fonts retain their SIL Open Font Licenses. Research informs the boundaries; no third-party corpus or transliteration implementation was copied.

Future extensions can add spell-checking, OCR suggestions, and better segmentation through separate modules. Those features are not included in v0.1.