Document Server@UHasselt >
Research >
Research publications >

Please use this identifier to cite or link to this item: http://hdl.handle.net/1942/743

Title: The exact rank-frequency function and size-frequency function of N-grams and N-word phrases with applications
Authors: EGGHE, Leo
Issue Date: 2005
Publisher: ELSEVIER
Abstract: N-grams are generalized words consisting of N consecutive symbols (letters), as they are used in a text. N-word phrases are general concepts consisting of N consecutive words, also as used in a text. Given the rank-frequency function of single letters (i.e. 1-grams) or of single words (i.e. 1-word phrases) being Zipfian, we determine in this paper the exact rank-frequency function (i.e. the occurrence of N-grams or N-word phrases on each rank) and size-frequency distribution (i.e. the density of N-grams or N-word phrases on each occurrence density) of these N-grams and N-word phrases. This paper distinguishes itself from other ones on this topic by allowing no approximations in the calculations. This leads to an intricate rank-frequency function for N-grams and N-word phrases (as we knew before from unpublished calculations) but leads surprisingly, to a very simple size-frequency function f(N) for N-grams or N-word phrases.
URI: http://hdl.handle.net/1942/743
DOI: 10.1016/j.mcm.2003.12.016
ISI #: 000229364100015
ISSN: 0895-7177
Category: A1
Type: Journal Contribution
Validation: ecoom, 2006
Appears in Collections: Research publications

Files in This Item:

Description SizeFormat
Published version824.36 kBAdobe PDF
Peer-reviewed author version482.65 kBAdobe PDF

Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.