Skip to content
View fauxneticien's full-sized avatar

Block or report fauxneticien

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
fauxneticien/README.md

Nay San

Staff Engineer at rime, working on data and modelling for conversational voice AI. Previously a PhD in Linguistics at Stanford (advised by Dan Jurafsky), on improving access to untranscribed speech corpora with AI.

More at resunay.com · Google Scholar


Speech: search and low-resource ASR

  • qbe-std_feats_eval — evaluation of feature-extraction methods for query-by-example spoken term detection in low-resource languages
  • bnf_cnn_qbe-std — query-by-example spoken term detection using bottleneck features and a CNN
  • u2u-asr — "user-to-user" ASR: a proof-of-concept workflow taking users from their own data to a fine-tuned model they can run locally in the browser (a play on "end-to-end ASR")
  • active_learning-w2v2_asr — active learning for fine-tuning wav2vec 2.0 ASR
  • asr-dataset-prep — scripts for preparing datasets for automatic speech recognition

Cross-lingual and self-supervised speech models

Phonetics and phonology

  • phonpack — an R package of fun(ctions) for doing phonetics
  • kphon — helper functions for the Kaytetye Phonological project (KPHON)
  • kaytetye-medial-vowels — processing scripts and datasets for a study of medial vowels in Kaytetye
  • akwelye — text-setting in akwelye (Kaytetye song)
  • census-languages — analysis of ABS Census data on Australian Indigenous languages

Lexicography and dictionaries

  • lexloop — iterative correction tool for data in a domain-specific language: edit a file, re-run, see validation and parsed views in the browser
  • lexicon-grammars — a collection of grammars for parsing backslash-coded lexicons
  • LexDev — a toolkit for generating live feedback on lexicographical data
  • kdict — data-processing functions for the Kaytetye Dictionary Transcriptions project
  • anamR — helper functions to read/write/process data from the Kaytetye database (KDB)

Pinned Loading

  1. CoEDL/vad-sli-asr CoEDL/vad-sli-asr Public

    A pipeline to isolate and transcribe one language in mixed-language speech

    Python 20 3

  2. qbe-std_feats_eval qbe-std_feats_eval Public

    Evaluation of feature extraction methods for query-by-example spoken term detection with low resource languages

    Perl 12 2

  3. CoEDL/vyov CoEDL/vyov Public

    Visualise your own vowels: A short introduction to Praat for complete beginners

    HTML 2

  4. CoEDL/tidylex CoEDL/tidylex Public

    Tidy lexicographical data in backslash-coded formats

    JavaScript 5

  5. CoEDL/yinarlingi CoEDL/yinarlingi Public

    R package for testing Warlpiri dictionary data structures

    R 1