Skip to content

Language Shootout

A user-supported site Fastest, Shortest, Simplest

K-Nucleotide

Python CPython #1 — K-Nucleotide

Count every k-nucleotide in a DNA sequence using the language’s own hash table, and report their frequencies.

Time 40,095.3 ms
CPU time 140,452.6 ms
Peak memory 380,532 KB
gz 612 bytes — comments removed, gzipped
Style ★★★★☆
Implementation Python — Python 3 (CPython) 3.13.7
Permitted dependency numpy — pinned in the toolchain image
By sysop-
Submitted September 24, 2026

Style assessment

A refined community submission that uses idiomatic Python where it counts: defaultdict(int), ProcessPoolExecutor parallelism, sorted with a tuple-reversal key, and join over generator expressions. Deductions for the global `sequence`, weak names (`l`, `r`, `frequences` misspelling), isinstance dispatch mixing int/str args in gen_result, and %-formatting; passing state explicitly, startswith(), and f-strings would raise it.

# The Computer Language Benchmarks Game
# http://benchmarksgame.alioth.debian.org/
#
# submitted by Ian Osgood
# modified by Sokolov Yura
# modified by bearophile
# modified by xfm for parallelization
# modified by Justin Peel 
# modified by Jean-Baptiste Lamy 

from sys import stdin
from collections import defaultdict

def gen_freq(frame):
    global sequence
    frequences = defaultdict(int)
    if frame == 1:
        for nucleo in sequence:
            frequences[nucleo] += 1
    else:
        for ii in range(len(sequence) - frame + 1) :
            frequences[sequence[ii : ii + frame]] += 1
    return frequences

def gen_result(arg):
    if isinstance(arg, int):
        frequences = gen_freq(arg)
        n = sum(frequences.values())
        l = sorted(frequences.items(), reverse = True, key = lambda seq_freq: seq_freq[::-1])
        return "".join("%s %.3f\n" % (st, 100.0 * fr / n) for st, fr in l) + "\n"
    else:
        frequences = gen_freq(len(arg))
        return "%s\t%s\n" % (frequences[arg], arg)

def prepare() :
    for line in stdin:
        if (line[0] == ">") and (line[1:3] == "TH"):
            break
        
    seq = ""
    for line in stdin:
        if line[0] == ">":
            break
        seq += line
    return seq.upper().replace('\n','')

def main():
    global sequence
    sequence = prepare()
    
    from concurrent.futures import ProcessPoolExecutor
    
    with ProcessPoolExecutor() as executor:
        r = executor.map(gen_result, ["GGTATTTTAATTTATAGT", "GGTATTTTAATT", "GGTATT", "GGTA", "GGT", 2, 1])
        
    print("".join(reversed(list(r))), end = "")
    
    
if __name__=='__main__' :
    main()

Back to K-Nucleotide