Wednesday, May 13, 2015

Malware and the Heavy Tail

Check out my new blog on Malware and the Heavy Tail over at the Verizon Security Blog!  Malware, (and many infosec feature), are unique in that they are long tailed.  The reality is that being long tailed means you have to treat them differently.  Find out how in the blog!

Thursday, May 7, 2015

The Circle of Life: A DBIR Attack Graph

Head on over to the Verizon security blog and check out my blog on turning the VERIS and the DBIR into an attack graph: The Circle of Life: A DBIR Attack Graph or on way back machine.  And keep an eye out.  This is just the primer for the juicy stuff!

Thursday, April 30, 2015

Automated Analysis of Competing Hypotheses for Diagnostic Medical Decision Support

It is with great pleasure I publish a project, hopefully to the benefit of society.  With my co-author Kindall Deitman, I have published a Diagnostic Medical Decision Support System which implements analysis of competing hypotheses.  It does so in the form of two new machine learning algorithms, the Bassett Deitmen Training Algorithm and the Bassett Deitmen Query Algorithm.  For those who follow this blog, it should come as no surprise that the underlying model is a directed graph.  For more information on the model, please reference the paper: Graph-based Diagnostic Medical Decision Support System.  You may watch our presentation of the project, where we answer a few questions as well, here.  If you would like to test or contribute to the model, it is licensed for non-commercial use here.  Should the need arise for a commercial license, please contact myself or Ms. Deitmen.

It is truly our hope that this model can benefit society as a whole.  Medical diagnosis is something that happens, regardless of the skill of the practitioner.  Whether it is a doctor with decade of experience, a new physician's assistant, a nurse in a small town or village, or even a bystander with little more than first aid training, when medical diagnosis is needed, it does not wait.  This project hopefully puts the experience of innumerable medical practitioners at the fingertips of those who need it most.  Additional, it hopefully brings a means of 'jogging the memory' of experienced medical practitioners who may have trouble remembering the obscure illnesses they rarely see.

There is certainly room for improvement.  A rather long list of updates and improvements already exists as we look to increase the utility of the project.  Nor is the approach constrained to medical diagnosis.  The training and query algorithms can support any situation in which observations or signals must be used to prioritize hypotheses as to the cause.  Still, there is a real and present need for such tools to help medical practitioners.  In an age where EMRs hold an incredible wealth of knowledge, it is our hope that this project may allow us to truly begin to harness it for the good of all.

0’day Campaigns for Everyone!

Hop over to the Verizon security blog to read my most recent O'day Campaigns for Everyone! (or why every attack now a'days looks targeted).  It's amazing what you can do with the DBIR data!

Friday, February 20, 2015

Association Rules

Check out my blog on Association Rules over at the Verizon Security Blog.  It's definitely an under-utilized area of machine learning.

Tuesday, January 6, 2015

Standardized Data Trees (UN M.49, ISO 3166-1, ISO 3366-2, country population/area, NAICS)

To help in aggregating, comparing, and validating data, I've created a pair of data of graphs which represent tree hierarchies of standard formatted data.

The World Graph contains:

World Graph Visualized


The NAICS Graph contains:
The percentage under the graph in the NAICS graph and the aggregate population/geographic area allow two things:
  1. Provide an amount as a dimension for data coded in these systems, (whether that amount be percentage of NAICS codes, population, or geographic area).
  2. Provide a means of comparing the similarity of two records by finding the Lowest Common Ancestor (LCA) and retrieving the score from that node.  The greater the score, the greater the distance between the nodes.
The hierarchies can be used for validating data as well as comparing things (as above) that are not on the same level of the graph.  For example, a US state could be compared to the country to South-Eastern Asia.

None of this is groundbreaking, but hopefully some find the graphs of use.