Applying the analog of the birthday problem to the RNG seeds explains why the project was launching processes with the same seed. Suppose you seed each process with an unsigned 16-bit integer. That means there are 65,536 possible seeds. Now suppose you launch 1,000 processes. With 65 times as many possible seeds as processes, surely every process should get its own seed, right? Not at all. There’s a 99.95% chance that two processes will have the same seed.
Showing posts with label statistics reporting and graphics. Show all posts
Showing posts with label statistics reporting and graphics. Show all posts
13 April 2016
Rando
John D. Cook alerts us to a pitfall in seeding a random number generator (RNG) for multiple processes/threads (for instance, for a Monte Carlo simulation):
02 January 2015
Breaking out of the black box
Sylvia Tippmann offers a brief introduction for scientists to the programming language R and its ecosystem.
27 October 2014
Any color is a good choice, so long as it's black
Nicolas P. Rougier et al. offer "Ten Simple Rules for Better Figures." The one that I tend to forget: Captions Are Not Optional. And the TL;DR version of the paper is captured by the first two rules: Know Your Audience and Identify Your Message.
29 May 2011
Shades of gray
Marcin Kozak proposes a hybrid design for scatterplots, using the best ideas of Edward Tufte and William Cleveland. Kozak's improvement really shines when applied to multipanel plots.
31 March 2011
08 January 2009
Not System R
Ashlee Vance introduces the open-source statistical analysis programming environment R, one that challenges proprietary packages.
"R has really become the second language for people coming out of grad school now, and there's an amazing amount of code being written for it," said Max Kuhn, associate director of nonclinical statistics at Pfizer. "You can look on the SAS message boards and see there is a proportional downturn in traffic."
06 September 2008
Pretty pictures
Anne Eisenberg reports on the IBM Many Eyes project, a contributor-driven site for data visualization.
19 August 2008
Top box
I'm in the process of migrating articles from an internal wiki to a new platform (I much prefer the one we're leaving, and I prefer MediaWiki to both, but it's not my call). Anyway, I stumbled over an entry that I wrote a few months ago about top box and bottom box statistics. Top box analysis, as far as I can tell, is fairly popular in the market research industry. It's often used with measures of customer satisfaction. It's a simple tool, but it hasn't received a lot of rigorous academic attention, so there isn't a lot of information available online. And, unfortunately, the only way to search for it is with "top box -office -set".
Subscribe to:
Posts (Atom)