Changelog¶
Note
Entries marked “via Claude” were developed with Claude Code. Dan directed the design, reviewed all output, and takes responsibility for the result. Unmarked entries by Dan were written without AI assistance.
Unreleased¶
Features:
Added
-m/--mime-typeflag tochardetectCLI — shows the detected MIME type alongside the encoding, similar tofile -i(Dan Blanchard via Claude, #384)
Accuracy:
cp864 Arabic now trains on contextually shaped text in visual order. The old training text forced every letter to its isolated presentation form in logical order, a byte pattern no real cp864 system ever wrote. (Dan Blanchard via Claude)
Visual-order Hebrew is now detected. The
he/iso8859-8model trains on both bidi conventions, so 1990s visual-order pages are detected as ISO-8859-8 while logical text prefers Windows-1255, the name signaling the storage convention as in Mozilla’s universal charset detector algorithm. (Dan Blanchard via Claude)
Performance:
Large inputs no longer need
max_bytesto stay fast.detect(data, max_bytes=len(data))on a 272 MiB buffer now finishes in about 0.13-0.23s, down from multiple seconds, on the compiled and pure-Python builds alike. UTF-8 validation is decode-based (CPython’s strict decoder enforces the identical rules at C speed, chunked so no matchingstris ever allocated) and stays exact over the whole window: the UTF-8 verdict is validated, never sampled. Candidate filtering, structural probing, the UTF-7/HZ deep validators, and rank corrections converge on the first 256 KB, a documented evidence cap that sits above the defaultmax_bytesso results for default calls are bit-for-bit unchanged (verified against the full corpus per file). Escape sequences are bounded by where they begin, not where they end, so one straddling that boundary is still judged whole. The new large-input tables in the performance docs cover the details; ADR-0006 in the repo records the design. (Dan Blanchard via Claude)detect(data, max_bytes=n)with n past that 256 KB cap again returns an encoding that decodes the whole window. Validity filtering converges on the cap, so a candidate that could not decode the rest of the window was able to reach the top: Windows-1252 for a 300 KB Latin-1 file whose only C1 bytes sit past the cap. The ranking’s winner is now decoded over the whole window before it is returned, and when it fails the next candidate that decodes takes over, the outcome 7.6.0 produced when validity read every byte. Default calls are unchanged; a window past the cap pays one C-speed decode of the winner, which puts 272 MiB of cp1252 at about 0.23s and of Shift_JIS at about 0.56s on the compiled build. (Dan Blanchard via Claude)ASCII and binary scans on large buffers no longer allocate input-sized transients:
detect_asciigates onisascii()and counts rather than materializing its remainder, andis_binarycounts through bounded chunks and derives its second control-byte total arithmetically instead of rescanning. (Dan Blanchard via Claude)The UTF-7
+scan no longer goes quadratic on line-wrapped base64. Its “is this+inside a base64 stream” guard skips newlines while walking backward, so on PEM certificates, MIME attachments, and JWTs every+rescanned everything before it. It now stops at the fourth character, which is all the verdict needs: 128 KiB of such data went from 5.4 seconds to 1.2 milliseconds. (Dan Blanchard via Claude)
Bug Fixes:
unittest.mock.patch.objecton the deprecatedchardet.equivalencesshim now restores the patched name on exit. The shim proxied reads and writes to the owning module but not deletes, so the mock raised on exit and leftchardet.evaluationpatched for the rest of the process. Deletes are proxied too, and ownership is fixed at import so the delete-then-restore round trip lands on the owner both times. (Dan Blanchard via Claude)prefer_superset=Trueno longer hands back a name that cannot decode the examined window. The Windows code pages leave a few C1 positions undefined that their ISO subsets map (0x81, 0x8D, 0x8F, 0x90, 0x9D under Windows-1252), and the remap applied regardless; it now checks the data and keeps the detected name when the superset cannot decode it.apply_preferred_supersettakes an optionaldataargument for the check. (Dan Blanchard via Claude)UniversalDetectorno longer reads a buffer that filled to exactlymax_bytesas the whole stream. Afeed()after the buffer was full did not mark it truncated, so the decode-safety flip could read a clipped multi-byte sequence as the end of the input and turn a correctutf-8answer into whichever single-byte codec decodes the fragment. (Dan Blanchard via Claude)detect_all()no longer resurrects a demoted niche Latin encoding. The demotion moved iso-8859-10, iso-8859-14, windows-1254, or HP-Roman8 to the end of the ranking but left it holding the top confidence, and the sortdetect_all()applies before returning put it back at second place. The demoted entry now takes the confidence of the candidate it sits behind, so the ranking the pipeline returns is ordered by confidence as documented. (Dan Blanchard via Claude)Language detection now honors
max_bytes. The language fill stage was handed the caller’s whole input and applied only its own 2 KB cap, sodetect(data, max_bytes=64)could still report a language derived from up to 2 KB. (Dan Blanchard via Claude)chardet.equivalencesis reachable again as an attribute of the package (import chardetthenchardet.equivalences.is_correct), which stopped working when__init__switched to importingchardet.output_names. The shim also resolves every name the pre-split module exposed, including the private tables and helpers; passes rebinding through to the module that owns the name, so patching a table on the shim reaches the code that reads it; and warns once per name used, naming that name and its new home, instead of once per process at import. Importing the shim is no longer an error under-W error::DeprecationWarning. (Dan Blanchard via Claude)A mostly-ASCII Windows-1252 file whose only non-ASCII content is a lone accented-letter pair (one
Ö/ö) no longer detects as HP-Roman8. The niche Latin demotion stood down whenever the data contained any byte the candidate decodes differently from ISO-8859-1, but presence is symmetric: Windows-1252 reads that 0xD6 asÖand HP-Roman8 asø, so both candidates contain the byte and it alone says nothing about which reading is right. The demotion now arbitrates the candidate against its swap target on the distinguishing bytes the way confusion resolution arbitrates its pairs. The bigram models decide when they know the bytes, so one Welshŵstill keeps ISO-8859-14. Word shape decides when they do not, so a Kvenđin Finnish prose keeps ISO-8859-10, a letter between letters beating the superscript Windows-1252 reads there. An evidence-free tie goes to the prevalent encoding. A candidate leading its swap target by more than the confusion band is not second-guessed on a handful of bytes. The demotion’s replacement is likewise no longer picked by sub-epsilon noise: among common Latin candidates tied with the best-placed of them, era prevalence chooses (Windows-1252 over ISO-8859-1), while a candidate leading its rivals by a real margin is still promoted on confidence. Some corpus files that came back as ISO-8859-1 now come back as Windows-1252, an accepted superset answer for them. (António Afonso via Claude, #383)An English file whose only non-ASCII letter is a capital
É,ÀorÓno longer detects as MacRoman. The MacRoman model reads those bytes as the ellipsis and curly quotes English text is full of, which handed it a lead worth 2e-5 and kept the era-prevalence prior from stepping in: the prior stood down whenever the leading model weighted any observed high-byte bigram. It now arbitrates such a leader against the prevalent candidate tied with it on the bytes the two read differently, every language variant on both sides, and promotes the prevalent candidate when it wins outright. A tie keeps the leader, so ISO-8859-1 results tied with Windows-1252 on data without C1 bytes are unchanged. (Dan Blanchard via Claude, #383)
Build:
scripts/verify_no_overlap.py --require-provenanceexits non-zero unless the training metadata records an exclusion set that matches today’s test data, for the release check. The shipped metadata currently records none: the iso8859-8 and cp864 subset retrains were built against later test-data states than the other 351 models, and only a full retrain can restore a single record for the set. (Dan Blanchard via Claude)Test coverage is back to 100% and CI now enforces it (
fail_under = 100). The newly covered paths include the EBCDIC plausibility checks in binary detection, charset declarations inside EBCDIC-encoded markup, the packed-kernel scoring branch (pinned bit-identical to the dense path), the confusion stage’s decisive-vote and strict-tier promotions, the classic-Mac and dead-heat rank corrections, and the model-file fallback/truncation paths. Two provably unreachable defensive guards are excluded with explanatorypragma: no covercomments instead of synthetic tests. (Dan Blanchard via Claude)The
models.bin/rowmax.bin/idf.binformats now have a single owner,chardet.models._format, called by both the trainer and the runtime loader; the row-maxima and IDF payloads regenerate byte-identically frommodels.bin, whose own bytes are guaranteed at the decompressed level (zlib output varies across builds). (Dan Blanchard via Claude)pipeline/markup.pyis now compiled by mypyc, bringing the list to 15 modules. The markup superset promotion moved into this module out of the compiled orchestrator, which left it running interpreted in every compiled wheel. Its structural comparison now scores the first 4 KB rather than the whole input, matching the window the charset declaration itself is validated on; the decode checks still read everything, since what a caller can decode is a fact about their whole input. (Dan Blanchard via Claude)
7.6.0 (2026-08-14)¶
Performance:
Compiled wheels now score bigram profiles through a small Cython kernel alongside mypyc, and the pair is 4.7x faster than the pure wheel on CPython 3.14.
_kernel.pystays plain Python (PyPy and pure wheels run it interpreted, unchanged),_kernel.pxdadds C types at build time and ships nothing, and detection output is bit-identical. The kernel declares itself safe without the GIL, so free-threaded CPython scales instead of silently re-enabling the GIL on import: 3.14t runs the whole suite in ~340ms across 8 threads, the fastest configuration measured. Compiled builds now need both hooks:HATCH_BUILD_HOOK_ENABLE_MYPYC=true HATCH_BUILD_HOOK_ENABLE_CUSTOM=true uv build
(Dan Blanchard via Claude)
Added support for CPython 3.15, including the free-threaded build. No code changes were needed. (Dan Blanchard via Claude)
Bug Fixes:
Fixed delimited ASCII data like
|NAME,+LAY|misdetecting as UTF-7, a follow-up to #371. Two new checks: the whole buffer must actually decode as UTF-7 (+|is an illegal shift, so tabular data fails immediately), and a block encoding a single code unit must land in a script range where a lone shifted character plausibly occurs.+LAYdecodes to U+2C06, Glagolitic; no genuine lone block in the corpus lands anywhere like it, while em dashes, ellipses, kanji, and accented letters all pass. (Dan Blanchard via Claude)Signed UTF-7 no longer reads as ASCII. The BOM stage recognizes the four UTF-7 signature prefixes (
+/v8-and friends) when the rest of the buffer decodes as UTF-7 — the prefix alone is ordinary ASCII (a diff of V8 source paths starts with+/v8). This is a deliberate divergence from WHATWG’s browser-security exclusion of UTF-7: chardet already detects the unsigned form, so refusing only the signed one made no sense. (Dan Blanchard via Claude)detect()no longer returns an encoding that cannot decode the input it was given (#380). When the whole input has been examined and the winner’s only multi-byte evidence is an incomplete trailing sequence, the best candidate that decodes the input completely wins instead. Genuinely truncated data keeps its answer: CJK cut mid-character, or input sliced atmax_bytes. (Dan Blanchard via Claude)Short apostrophe-heavy English is no longer labeled Scottish Gaelic or Breton. A rare-language label on an input under 128 bytes now needs a 0.03 lead over the best mainstream language, mirroring the encoding-side arbitration in ADR-0005. (Dan Blanchard via Claude)
Hungarian text no longer loses to a Czech reading: confusion rescoring compares tied pairs only under language models both encodings have. (Dan Blanchard via Claude)
Space-padded text no longer matches a degenerate Serbian model at high confidence. Statistical scoring now skips repeated-whitespace bigrams, matching the whitespace collapse training already applies. Fixed 21 files plus a long-standing GB2312 known failure. (Dan Blanchard via Claude)
EBCDIC text is no longer invisible to the early pipeline. The binary stage treats EBCDIC’s 0x05/0x15 tab and newline as whitespace when the data is high-byte-dominated, and the markup stage reads charset declarations through a cp037 decode of the head. (Dan Blanchard via Claude)
Fixed the last two EBCDIC sibling misdetections: a letter reading beating punctuation is no longer evidence by itself, and low-confidence near-ties scan deeper but need the category vote and the bigram rescore to agree. (Dan Blanchard via Claude)
Fixed training normalization gaps that starved ISO-8859-16 (legacy cedilla forms) and the 26 pre-euro encodings (the euro sign) at exactly the bytes that discriminate them from their siblings. Cedilla folds to comma-below and the euro to the currency sign wherever the target encoding cannot represent them. (Dan Blanchard via Claude)
Improvements:
Retrained every bigram model on a refreshed, deduplicated corpus with the ADR-0004 hardening: whitespace collapses after encoding, retention guards hard-fail mostly unencodable corpora, Serbian gets real Latin-script text, CP1006 gains its sixteen missing Urdu letters, and wiki markup is stripped before bigram counting. Training provenance is now recorded per model, so test data added after a retrain is detectable as such. (Dan Blanchard via Claude)
New ANSI-art model. cp437 detection now includes a profile trained on 16,621 text-mode art files from 16colo.rs, keyed under the
zxxpseudo-language and reported withlanguage=None. (Dan Blanchard via Claude)Rare-language arbitration: a low-confidence statistical winner from a language with no documented legacy-encoding population (Scottish Gaelic, Welsh, Irish, Breton) yields to a near-tied mainstream candidate. Genuine Celtic text wins by landslides and is unaffected; eight boundary sentinels in the test suite guard the gate. Design and evidence in ADR-0005. (Dan Blanchard via Claude)
Confusion-group resolution is context-aware: votes are counted per occurrence, letter readings with no word shape are demoted, and art-model wins are exempt. Fixed twelve EBCDIC and Latin files that were riding single-byte coin flips. (Dan Blanchard via Claude)
Statistical dead heats no longer resolve by candidate enumeration order. Three tiebreaks: prefer the Windows superset, prefer the more prevalent era when there is no high-byte evidence, and prefer a classic-Mac candidate when line endings are bare
\r. (Dan Blanchard via Claude)Training pipeline hardening after a cache-loss post-mortem: retrains that would silently drop a model now abort loudly, caches filter in place instead of being deleted wholesale, and the artpack fetcher builds into a temporary directory. (Dan Blanchard via Claude)
7.5.1 (2026-08-06)¶
Bug Fixes:
Fixed markup-declared encodings being reported under a name that can’t decode the input. A page declaring
Shift_JISbut using CP932 extension characters (like ①) came back asSHIFT_JIS, which fails.decode()on those same bytes. Superset promotion (CP932,CP949) now always fires when the reported name can’t decode the data but the superset can. (Dan Blanchard via Claude)Fixed a lying charset declaration beating genuine UTF-8 content. A UTF-8 page declaring
<meta charset="iso-8859-1">came back as ISO-8859-1, which decodes to mojibake. Valid multi-byte UTF-8 now wins over a conflicting declaration; pure ASCII and real single-byte content still honor it. (Dan Blanchard via Claude)Fixed BOM-less UTF-16 byte-order detection for pure-CJK text. With no ASCII in the sample, the only null bytes come from the low byte of characters like U+4E00 (一), which sit in the wrong parity position, so short Chinese UTF-16 samples came back with reversed endianness at full confidence. Byte order is now chosen by decoding both ways and comparing text quality, with the null signal breaking near-ties. Found by scoring chardet against charset-normalizer’s char-dataset. (Dan Blanchard via Claude)
7.5.0 (2026-08-05)¶
Bug Fixes:
Fixed multi-byte encodings being eliminated when the input ends in an incomplete character. Byte-validity filtering decoded with a one-shot strict decode, which cannot tell a truncated tail from corrupt data, so a single dangling lead byte dropped every CJK candidate and the result came down to input-length parity — a 184-byte GBK sample detected as
GB18030, the same sample minus one byte asWindows-1256. This was also reachable on complete, well-formed files, because chardet slices its own input atmax_bytesand at_SCAN_LIMITin_validate_bytes(): a valid 14 kB GBK page with an honest<meta charset="gbk">lost its declaration, and with ittext/htmland 0.95 confidence, whenever byte 4096 happened to split a character. Validity checks now decode incrementally withfinal=False, deferring a partial trailing character while still rejecting corruption anywhere before it. (António Afonso via Claude, #376)Fixed
compat_names(the default) leaking internal Python codec names for seven encodings.detect()now returnsISO-8859-2,ISO-8859-6,ISO-8859-13,Windows-1250,Windows-1256,Windows-1257, andCP874instead of their lowercase codec spellings. These were absent from_COMPAT_NAMESafter the 7.1.0 switch to codec-name canonicals, which made default output inconsistent with their siblings (e.g.cp1250vsWindows-1251) and with the encoding-name table in Usage. (António Afonso via Claude, #374)Fixed
compat_names(the default) leaking the internalcp932codec name.detect()now returnsCP932instead ofcp932, matching its Japanese siblings (shift_jis_2004→SHIFT_JIS) and the value chardet 5.x/6.x returned. (uttam12331, #375)
Performance:
Statistical scoring now skips single-byte models that provably can’t beat the current runner-up, using per-model row-maximum tables (
rowmax.bin). Results are bit-identical: multi-byte models are always scored in full anddetect_all()bypasses pruning. Mean detection time dropped ~2.9x with mypyc, with the largest gains on legacy CJK (p99 from 9.5ms to 3.3ms). (Dan Blanchard via Claude)Model tables are now
bytes(native array indexing under mypyc, instead of boxedmemoryviewcalls), and the model blob is decompressed in chunks rather than one shot: peak process memory dropped from 53.9 to 27.4 MiB. (Dan Blanchard via Claude)Confusion-group resolution and post-processing now use
bytes.translateprefilters instead of per-byte Python scans, making near-tie resolution cheaper on large inputs. (Dan Blanchard via Claude)
Improvements:
prefer_superset=Trueis now documented as the recommended mode and will become the default in chardet 8.0. Detection examines at mostmax_bytesof input, so only the superset encoding is guaranteed to decode bytes beyond that window — the same reasoning behind the WHATWG/W3C Encoding Standard’s rule that browsers decodeasciiandiso-8859-1content aswindows-1252. Callers that depend on subset names should start passingprefer_superset=Falseexplicitly. (Dan Blanchard via Claude)chardet.equivalencesis now a deprecation shim. Accuracy-evaluation predicates (is_correct,is_equivalent_detection, etc.) moved tochardet.evaluation; public-API encoding-name remapping (apply_compat_names,apply_preferred_superset) moved tochardet.output_names. Existing imports keep working with aDeprecationWarning.chardet.equivalenceswill be removed in 8.0. (Dan Blanchard via Claude)Internal pipeline reorganization: language detection, markup-superset promotion, and post-processing rank corrections moved out of the orchestrator into
pipeline/language.py,pipeline/markup.py, andpipeline/postprocess.pyrespectively. No behavior change. The two new modules are also added to the mypyc compilation list. (Dan Blanchard via Claude)
7.4.3 (2026-04-13)¶
Bug Fixes:
Fixed
ValueError: embedded null charactercrash when input contained a<meta charset>declaration with a null byte in the encoding name (e.g.b'<meta charset="\x00utf-8">').codecs.lookup()raisesValueErroron embedded nulls, andlookup_encoding()was only catchingLookupError. Also added defensiveValueErrorcatches in_validate_bytes()and_to_utf8()for completeness. (Dan Blanchard via Claude, #369)
7.4.2 (2026-04-12)¶
Bug Fixes:
Fixed
RuntimeError: pipeline must always return at least one resulton ~2% of all possible two-byte inputs (e.g.b"\xf9\x92"). Multi-byte encodings like CP932 and Johab could score above the structural confidence threshold on very short inputs, but then statistical scoring would return nothing, leaving the pipeline with an empty result list instead of falling through to theno_match_encodingfallback. (Jason Barnett via Claude, #367, #368)
Improvements:
Added ~90 encoding aliases from the WHATWG Encoding Standard and IANA Character Sets registry so that
<meta charset>labels likex-cp1252,x-sjis,dos-874,csUTF8, and thecswindows*family all resolve correctly through the markup detection stage. Every alias was driven by a failing spec-compliance test. (Dan Blanchard via Claude, #366)Added a spec-compliance test suite covering Python decode round-trips for all 86 registry encodings, WHATWG web-platform label resolution, IANA preferred MIME names, and Unicode/RFC conformance (BOM sniffing, UTF-8 boundary cases, UTF-16 surrogate pairs). This is the test suite that would have caught the 7.4.1 BOM bug before release. (Dan Blanchard via Claude, #366)
7.4.1 (2026-04-07)¶
Bug Fixes:
BOM-prefixed UTF-16 and UTF-32 input now reports
utf-16andutf-32instead of the endian-specific variants. Python’sutf-16-le/utf-16-be/utf-32-le/utf-32-becodecs keep the BOM as a U+FEFF in the decoded string, whileutf-16/utf-32strip it, so callers passing the detection result directly to.decode()were getting a stray BOM at the start of their text. BOM-less UTF-16/32 detection (via null-byte patterns) is unchanged and still returns the endian-specific name. (Dan Blanchard via Claude, #364, #365)
7.4.0 (2026-03-26)¶
Performance:
Switched to dense zlib-compressed model format (v2): models are now stored as contiguous
memoryviewslices of a single decompressed blob, eliminating per-modelstruct.unpackoverhead. Cold start (import + first detect) dropped from ~75ms to ~13ms with mypyc. (Dan Blanchard via Claude, #354)
Accuracy:
Accuracy improved from 98.6% to 99.3% (2499/2517 files) through a combination of training and scoring improvements:
Eliminated train/test data overlap by content-fingerprinting test suite articles and excluding them from training data (#351)
Added MADLAD-400 and Wikipedia as supplemental training sources to fill gaps left by exclusion filtering (#351)
Improved non-ASCII bigram scoring: high-byte bigrams are now preserved during training (instead of being crushed by global normalization), and weighted by per-bigram IDF so encoding-specific byte patterns contribute proportionally to how discriminative they are (#352)
Added encoding-aware substitution filtering: character substitutions during training now only apply for characters the target encoding cannot represent
Increased training samples from 15K to 25K per language/encoding pair (Dan Blanchard via Claude)
Bug Fixes:
Added dedicated structural analyzers for CP932, CP949, and Big5-HKSCS: these superset encodings previously shared their base encoding’s byte-range analyzer, missing extended ranges unique to each superset (Dan Blanchard via Claude, #353)
7.3.0 (2026-03-24)¶
License:
0BSD license — the project license has been changed from MIT to 0BSD, a maximally permissive license with no attribution requirement. All prior 7.x releases should also be considered 0BSD licensed as of this release. (Dan Blanchard via Claude)
Features:
Added
mime_typefield to detection results — identifies file types for both binary (via magic number matching) and text content. Returned in alldetect(),detect_all(), andUniversalDetectorresults. (Dan Blanchard via Claude, #350)New
pipeline/magic.pymodule detects 40+ binary file formats including images, audio/video, archives, documents, executables, and fonts. ZIP-based formats (XLSX, DOCX, JAR, APK, EPUB, wheel, OpenDocument) are distinguished by entry filenames. (Dan Blanchard via Claude, #350)
Bug Fixes:
Fixed incorrect equivalence between UTF-16-LE and UTF-16-BE in accuracy testing — these are distinct encodings with different byte order, not interchangeable (Dan Blanchard via Claude)
Performance:
Added 4 new modules to mypyc compilation (orchestrator, confusion, magic, ascii), bringing the total to 11 compiled modules (Dan Blanchard via Claude)
Capped statistical scoring at 16 KB — bigram models converge quickly, so large files no longer score the full 200 KB. Worst-case detection time dropped from 62ms to 26ms with no accuracy loss. (Dan Blanchard via Claude)
Replaced
dataclasses.replace()with directDetectionResultconstruction on hot paths, eliminating ~354k function calls per full test suite run (Dan Blanchard via Claude)
Build:
Added riscv64 to the mypyc wheel build matrix — prebuilt wheels are now published for RISC-V Linux alongside existing architectures (Bruno Verachten, #348)
7.2.0 (2026-03-17)¶
Features:
Added
include_encodingsandexclude_encodingsparameters todetect(),detect_all(), andUniversalDetector— restrict or exclude specific encodings from the candidate set, with corresponding-i/--include-encodingsand-x/--exclude-encodingsCLI flags (Dan Blanchard via Claude, #343)Added
no_match_encoding(default"cp1252") andempty_input_encoding(default"utf-8") parameters — control which encoding is returned when no candidate survives the pipeline or the input is empty, with corresponding CLI flags (Dan Blanchard via Claude, #343)Added
-l/--languageflag tochardetectCLI — shows the detected language (ISO 639-1 code and English name) alongside the encoding (Dan Blanchard via Claude, #342)
7.1.0 (2026-03-11)¶
Features:
Added PEP 263 encoding declaration detection —
# -*- coding: ... -*-and# coding=...declarations on lines 1–2 of Python source files are now recognized with confidence 0.95 (Dan Blanchard via Claude, #249)Added
chardet.universaldetectorbackward-compatibility stub so thatfrom chardet.universaldetector import UniversalDetectorworks with a deprecation warning (Dan Blanchard via Claude, #341)
Fixes:
Fixed false UTF-7 detection of ASCII text containing
++or+wordpatterns (Dan Blanchard, #332, #335)Fixed 0.5s startup cost on first
detect()call — model norms are now computed during loading instead of lazily iterating 21M entries (Dan Blanchard via Claude, #333, #336)Fixed undocumented encoding name changes between chardet 5.x and 7.0 —
detect()now returns chardet 5.x-compatible names by default (Dan Blanchard via Claude, #338)Improved ISO-2022-JP family detection — recognizes ESC sequences for ISO-2022-JP-2004 (JIS X 0213) and ISO-2022-JP-EXT (JIS X 0201 Kana) (Dan Blanchard via Claude)
Fixed silent truncation of corrupt model data (
iter_unpackyielded fewer tuples instead of raising) (Dan Blanchard via Claude)Fixed incorrect date in LICENSE (Dan Blanchard)
Performance:
5.5x faster first-detect time (~0.42s → ~0.075s) by computing model norms as a side-product of
load_models()(Dan Blanchard via Claude)~40% faster model parsing via
struct.iter_unpackfor bulk entry extraction (eliminates ~305K individualunpackcalls) (Dan Blanchard via Claude)
New API parameters:
Added
compat_namesparameter (defaultTrue) todetect(),detect_all(), andUniversalDetector— set toFalseto get raw Python codec names instead of chardet 5.x/6.x compatible display names (Dan Blanchard via Claude)Added
prefer_supersetparameter (defaultFalse) — remaps legacy ISO/subset encodings to their modern Windows/CP superset equivalents (e.g., ASCII → Windows-1252, ISO-8859-1 → Windows-1252). This will default to ``True`` in the next major version (8.0). (Dan Blanchard via Claude)Deprecated
should_rename_legacyin favor ofprefer_superset— a deprecation warning is emitted when used (Dan Blanchard via Claude)
Improvements:
Switched internal canonical encoding names to Python codec names (e.g.,
"utf-8"instead of"UTF-8"), withcompat_namescontrolling the public output format. See Usage for the full mapping table. (Dan Blanchard via Claude)Added
lookup_encoding()toregistryfor case-insensitive resolution of arbitrary encoding name input to canonical names (Dan Blanchard via Claude)Achieved 100% line coverage across all source modules (+31 tests) (Dan Blanchard via Claude)
Updated benchmark numbers: 98.2% encoding accuracy, 95.2% language accuracy on 2,510 test files (Dan Blanchard via Claude)
Pinned test-data cloning to chardet release version tags for reproducible builds (Dan Blanchard via Claude)
7.0.1 (2026-03-04)¶
Fixes:
Fixed false UTF-7 detection of SHA-1 git hashes (Alex Rembish, #324)
Fixed
_SINGLE_LANG_MAPmissing aliases for single-language encoding lookup (e.g.,big5→big5hkscs) (Dan Blanchard)Fixed PyPy
TypeErrorin UTF-7 codec handling (Dan Blanchard)
Improvements:
Retrained bigram models — 24 previously failing test cases now pass (Dan Blanchard via Claude)
Updated language equivalences for mutual intelligibility (Slovak/Czech, East Slavic + Bulgarian, Malay/Indonesian, Scandinavian languages) (Dan Blanchard via Claude)
7.0.0 (2026-03-02)¶
Ground-up, 0BSD-licensed rewrite of chardet (Dan Blanchard via Claude, #322). Same package name, same public API — drop-in replacement for chardet 5.x/6.x.
Highlights:
0BSD license (previous versions were LGPL)
96.8% accuracy on 2,179 test files (+2.3pp vs chardet 6.0.0, +7.7pp vs charset-normalizer)
41x faster than chardet 6.0.0 with mypyc (28x pure Python), 7.5x faster than charset-normalizer
Language detection for every result (90.5% accuracy across 49 languages)
99 encodings across six eras (MODERN_WEB, LEGACY_ISO, LEGACY_MAC, LEGACY_REGIONAL, DOS, MAINFRAME)
12-stage detection pipeline — BOM, UTF-16/32 patterns, escape sequences, binary detection, markup charset, ASCII, UTF-8 validation, byte validity, CJK gating, structural probing, statistical scoring, post-processing; the markup stage’s PEP 263 declaration sniffing was requested by patrikha in #249
Bigram frequency models trained on CulturaX multilingual corpus data for all supported language/encoding pairs
Optional mypyc compilation — 1.49x additional speedup on CPython
Thread-safe
detect()anddetect_all()with no measurable overhead; scales on free-threaded Python 3.13t+Negligible import memory (96 B)
Zero runtime dependencies
Breaking changes vs 6.0.0:
detect()anddetect_all()now default toencoding_era=EncodingEra.ALL(6.0.0 defaulted toMODERN_WEB)Internal architecture is completely different (probers replaced by pipeline stages). Only the public API is preserved.
LanguageFilteris accepted but ignored (deprecation warning emitted)chunk_sizeis accepted but ignored (deprecation warning emitted)
6.0.0.post1 (2026-02-22)¶
Fixed
__version__not being set correctly in the package (Dan Blanchard)
6.0.0 (2026-02-22)¶
Features:
Unified single-byte charset detection with proper language-specific bigram models for all single-byte encodings (replaces
Latin1ProberandMacRomanProberheuristics) (Dan Blanchard)38 new languages: Arabic, Belarusian, Breton, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Farsi, Finnish, French, German, Icelandic, Indonesian, Irish, Italian, Kazakh, Latvian, Lithuanian, Macedonian, Malay, Maltese, Norwegian, Polish, Portuguese, Romanian, Scottish Gaelic, Serbian, Slovak, Slovene, Spanish, Swedish, Tajik, Ukrainian, Vietnamese, Welsh (Dan Blanchard)
EncodingErafiltering via newencoding_eraparameter (Dan Blanchard)max_bytesandchunk_sizeparameters fordetect(),detect_all(), andUniversalDetector; chunked processing was proposed by deedy5 in #284 (Dan Blanchard)-e/--encoding-eraCLI flag (Dan Blanchard via Claude)EBCDIC detection (CP037, CP500) (Dan Blanchard)
Direct GB18030 support (replaces redundant GB2312 prober) (Dan Blanchard)
Binary file detection (Dan Blanchard)
Python 3.12, 3.13, and 3.14 support (Hugo van Kemenade, #283)
GitHub Codespaces support (oxygen dioxide, #312)
Breaking changes:
Dropped Python 3.7, 3.8, and 3.9 (requires Python 3.10+)
Removed
Latin1ProberandMacRomanProberRemoved EUC-TW support
Removed
LanguageFilter.NONEdetect()default changed toencoding_era=EncodingEra.MODERN_WEB
Fixes:
Fixed SJIS distribution analysis (second-byte range >= 0x80) (Kadir Can Ozden, #315)
Fixed
max_bytesnot being passed toUniversalDetector(Kadir Can Ozden, #314)Fixed UTF-16/32 detection for non-ASCII-heavy text (Dan Blanchard)
Fixed GB18030
char_len_table(Dan Blanchard)Fixed UTF-8 state machine (Dan Blanchard)
Fixed
detect_all()returning inactive probers (Dan Blanchard)Fixed early cutoff bug (Dan Blanchard)
Updated LGPLv2.1 license text for remote-only FSF address (Ben Beasley, #307)
5.2.0 (2023-08-01)¶
Added support for running the CLI via
python -m chardet(Dan Blanchard)
5.1.0 (2022-12-01)¶
Added
should_rename_legacyargument to remap legacy encoding names to modern equivalents (Dan Blanchard, #264)Added MacRoman encoding prober (Elia Robyn Lake)
Added
--minimalflag tochardetectCLI (Dan Blanchard, #214)Added type annotations and mypy CI (Jon Dufresne, #261)
Added support for Python 3.11 (Hugo van Kemenade, #274)
Added ISO-8859-15 capital letter sharp S handling (Simon Waldherr, #222)
Clarified LGPL version in license trove classifier (Ben Beasley, #255)
Removed support for Python 3.6 (Jon Dufresne, #260)
5.0.0 (2022-06-25)¶
Added UTF-16/32 BE/LE probers (Jason Zavaglia, #109, #206)
Added test data for Croatian, Czech, Hungarian, Polish, Slovak, Slovene, Greek, Turkish (Dan Blanchard)
Improved XML tag filtering (Dan Blanchard, #208)
Made
detect_allreturn child prober confidences (Dan Blanchard, #210)Added support for Python 3.10 (Hugo van Kemenade, #232)
Dropped Python 2.7, 3.4, 3.5 (requires Python 3.6+)
4.0.0 (2020-12-10)¶
Added
detect_all()function returning all candidate encodings (Damien, #111)Converted single-byte charset probers to nested dicts (performance) (Dan Blanchard, #121)
CharsetGroupProbernow short-circuits on definite matches (performance) (Dan Blanchard, #203)Added
languagefield todetect_alloutput (Dan Blanchard)Switched from Travis to GitHub Actions (Dan Blanchard, #204)
Dropped Python 2.6, 3.4, 3.5
3.0.4 (2017-06-08)¶
Fixed packaging issue with
pytest_runner(Zac Medico, #119)Included
test.pyin source distribution (Zac Medico, #118)Updated old URLs in README and docs (Qi Fan, #123; Jon Dufresne, #129)
3.0.3 (2017-05-16)¶
Fixed crash when debug logging was enabled (Dan Blanchard, #117)
3.0.2 (2017-04-12)¶
Fixed
detectsometimes returningNoneinstead of a result dict (Dan Blanchard, #114)
3.0.1 (2017-04-11)¶
Fixed crash in EUC-TW prober with certain strings (Dan Blanchard)
3.0.0 (2017-04-11)¶
Added Turkish ISO-8859-9 detection (queeup)
Modernized naming conventions (
typical_positive_ratioinstead ofmTypicalPositiveRatio) (Dan Blanchard, #107)Added
languageproperty to probers and results (Dan Blanchard, #108)Switched from Travis to GitHub Actions (Dan Blanchard)
Fixed
CharsetGroupProber.statenot being set toFOUND_IT(Dan Blanchard)Added Hypothesis-based fuzz testing (David R. MacIver, #66)
Don’t indicate byte order for UTF-16/32 with given BOM, for compatibility with
decode()(Sebastian Noack, #73)Stop reading file immediately when file type is known (Jason Zavaglia, #103)
chardet 2.3.0 (2014-10-07)¶
Added CP932 detection (hashy)
Switched
chardetectto useargparse(Dan Blanchard)
chardet 2.2.1 (2013-12-18)¶
chardet 2.2.0 (2013-12-16)¶
Merged the charade fork back into chardet, unifying Python 2 and Python 3 support under the original package name.
Added CP949 detection (Kyung-hown Chung)
Fixed BOM detection (Jean Boussier)
charade 1.0.3 (2013-01-18)¶
Fixed codecs usage for compatibility (Ian Cordasco)
charade 1.0.2 (2013-01-18)¶
Fixed BOM detection (Jean Boussier)
Improved multibyte sequence handling (Kyung-hown Chung)
charade 1.0.1 (2012-12-03)¶
Version fix (Ian Cordasco)
charade 1.0.0 (2012-12-02)¶
Initial release: Python 3 port of chardet, forked as a separate package (Ian Cordasco)
chardet 2.1.1 (2012-10-01)¶
Bumped version past Mark Pilgrim’s last release
chardetectcan now read from stdin (Erik Rose)Fixed BOM byte strings for UCS-4-2143 and UCS-4-3412 (Toshio Kuratomi)
Restored Mark Pilgrim’s original docs and COPYING file (Toshio Kuratomi)
chardet 1.1 (2012-07-27)¶
Added
chardetectCLI tool (Erik Rose)Fixed
utf8probercrash when character is out of range (David Cramer)Cleaned up detection logic to fail gracefully (David Cramer)
Fixed feed encoding errors (David Cramer)
chardet 1.0.1 (2008-04-19)¶
Packaging fix, added egg distributions for Python 2.4 and 2.5 (Mark Pilgrim)
chardet 1.0 (2006-12-23)¶
Initial release: Python 2 port of Mozilla’s universal charset detector (Mark Pilgrim)