Fasta
Variance
Please don't optimize the cumulative-probabilities lookup (for example, by using a scaling factor) or naïve LCG arithmetic - those programs will not be accepted.
How to implement
We ask that contributed programs not only give the correct result, but also use the same algorithm to calculate that result.
Each program should:
- Generate DNA sequences, by copying from a given sequence
- Generate DNA sequences, by weighted random selection from 2 alphabets:
- Convert the expected probability of selecting each nucleotide into cumulative probabilities
- Match a random number against those cumulative probabilities to select each nucleotide (use linear search or binary search)
- Use this naïve linear congruential generator to calculate a random number each time a nucleotide needs to be selected (don't cache the random number sequence):
IM = 139968 IA = 3877 IC = 29573 Seed = 42 Random(Max) Seed = (Seed * IA + IC) modulo IM = Max * Seed / IM
Verification: Use diff to compare program output N=1000 with the reference output.
Use a larger command line argument (25000000) to check program performance.
Times are wall-clock milliseconds, with this implementation’s hello-world startup time subtracted. gz is the source in bytes with comments removed and gzipped. style is the idiomatic-code score. Click a heading to sort.