Identifies the dominant language of a text sample by asking an LLM (OpenAI
or Gemini) to pick one name from a candidate list. More accurate than
detect_language() on short texts, mixed-language corpora, or languages
with sparse snowball stopword coverage, at the cost of an API call.
Usage
detect_language_llm(
texts,
languages,
provider = c("auto", "openai", "gemini"),
model = NULL,
api_key = NULL,
sample_n = 200,
seed = 123,
verbose = TRUE
)Arguments
- texts
Character vector of documents.
- languages
Named character vector of candidate languages (name = display name, value = language code), e.g.
c(English = "en", French = "fr").- provider
One of
"auto","openai","gemini"."auto"picks whichever ofOPENAI_API_KEY/GEMINI_API_KEYis set, or the key implied byapi_key's prefix.- model
Model name (default depends on provider).
- api_key
API key; falls back to the provider's environment variable.
- sample_n
Maximum documents to sample for the prompt (default 200).
- seed
Seed for sampling.
- verbose
Logical, print status messages (default TRUE).
Value
A one-row tibble with language (code) and provider, or NULL
when no provider/key is available or the response doesn't match a
candidate.
See also
detect_language() for the local, no-API alternative; call_llm_api() for the direct provider call.
Examples
if (interactive()) {
detect_language_llm(
c("Bonjour, comment allez-vous?", "Je suis ravi de vous rencontrer."),
languages = c(English = "en", French = "fr", Spanish = "es"),
provider = "openai",
api_key = Sys.getenv("OPENAI_API_KEY")
)
}
