fix(locale): use only common surnames for zh_CN last names - #3938
Merged
Merged
Conversation
The zh_CN surname list held ~1000 entries, most of them ultra-rare surnames some Chinese speakers never encounter, drawn uniformly — so most generated names looked unnatural at a glance. Replace it with the curated list of the most common surnames (including compound surnames), which cover the vast majority of the population, mirroring the zh_TW list's scope. Closes #3466
✅ Deploy Preview for fakerjs ready!Built without sensitive environment variables
To edit notification comments on pull requests, go to your Netlify project configuration. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## next #3938 +/- ##
=======================================
Coverage 98.84% 98.84%
=======================================
Files 923 923
Lines 3216 3216
Branches 583 583
=======================================
Hits 3179 3179
Misses 33 33
Partials 4 4 🚀 New features to boost your workflow:
|
matthewmayer
approved these changes
Jul 5, 2026
ST-DDT
approved these changes
Jul 5, 2026
Shinigami92
approved these changes
Jul 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3466
Problem
zh_CNlast names are drawn uniformly from a ~1000-entry list in which the large majority are ultra-rare surnames (不, 丑, 九, 业, 之, …) that many Chinese speakers never encounter in real life. As the issue reporter showed, most generated names look unnatural at first glance because the surname is almost always one of these rare ones.Fix (data-only, per maintainer guidance)
@ST-DDT confirmed in the thread that "both weighting and/or reducing the number of rare last names in that list are valid solutions". With the weighting PoC (#3467) closed, this PR takes the reduction route: replace
last_name.genericwith the curated list of the most common surnames posted in the issue by @yyz945947732 — 150 single-character surnames plus 12 common compound surnames (欧阳, 诸葛, 司马, …), covering ~85%+ of the population. This mirrors the scope of the existingzh_TWlist (~100 common surnames). No code changes;last_name_patternis untouched.The file was normalized with
pnpm run generate:locales, and thelocale-datacharacter-inventory snapshot updated accordingly (one line:zh_CN).Verification
Added a regression test in
test/modules/person.spec.tsasserting thezh_CNlist is the curated common set (contains 王/李/张/刘/陈/欧阳, excludes the rare 不/丑/九/业/之, length < 200). It fails on the old data and passes with this change. Full runs green locally:person.spec(214),locale-data(2635),locale-imports(152),all-functional(36k),ts-check, lint.Credit for the curated list goes to @yyz945947732 in the issue.