{"id":2532,"date":"2010-06-01T06:51:44","date_gmt":"2010-06-01T06:51:44","guid":{"rendered":"http:\/\/www.smartdatacollective.com\/index.php\/post\/27599\/"},"modified":"2010-06-01T06:51:44","modified_gmt":"2010-06-01T06:51:44","slug":"27599","status":"publish","type":"post","link":"https:\/\/www.smartdatacollective.com\/27599\/","title":{"rendered":"Voodoo Spectrum of Machine Learning and Data Sets"},"content":{"rendered":"<p>I used to be very gung-ho about machine learning approaches to trading but I&#8217;m less so now. You have to understand that that there is a spectrum of alpha sources, from very specific structured arbitrage opportunities -&gt; to stat arb -&gt; to just voodoo nonsense. <\/p>\n<div> <\/div>\n<div>As history goes on, hedge funds and other large players are absorbing the alpha from left to right. Having squeezed the pure arbs (ADR vs underlying, ETF vs components, mergers, currency triangles, etc) they then became hungry again and moved to stat arb (momentum, correlated pairs, regression analysis, news sentiment, etc). But now even the big stat arb strategies are running dry so people go further, chasing mirages (nonlinear regression, causality inference in large data sets, etc). <\/p>\n<div> <\/div>\n<div>In modeling the market, it&#8217;s best to start with as much structure as possible before moving on to more amorphous statistical strategies. If you have to use statistical machine learning, encode as much trading domain knowledge as possible with specific distance\/neighborhood metrics, linearity, variable importance weightings, hierarchy, low-dimensional factors, etc. <\/div>\n<div> <\/div>\n<div>It&#8217;s good to have a heuristic feel for the <span class=\"dots\">&#8230;<\/span><\/div>\n<\/p><\/div>\n<p><!--more--><\/p>\n<p><!--break--><br \/>\nI used to be very gung-ho about machine learning approaches to trading but I&#8217;m less so now. You have to understand that that there is a spectrum of alpha sources, from very specific structured arbitrage opportunities -&gt; to stat arb -&gt; to just voodoo nonsense.<\/p>\n<div>\n<\/div>\n<div>As history goes on, hedge funds and other large players are absorbing the alpha from left to right. Having squeezed the pure arbs (ADR vs underlying, ETF vs components, mergers, currency triangles, etc) they then became hungry again and moved to stat arb (momentum, correlated pairs, regression analysis, news sentiment, etc). But now even the big stat arb strategies are running dry so people go further, chasing mirages (nonlinear regression, causality inference in large data sets, etc).<\/p>\n<div>\n<\/div>\n<div>In modeling the market, it&#8217;s best to start with as much structure as possible before moving on to more amorphous statistical strategies. If you have to use statistical machine learning, encode as much trading domain knowledge as possible with specific distance\/neighborhood metrics, linearity, variable importance weightings, hierarchy, low-dimensional factors, etc. <\/div>\n<div>\n<\/div>\n<div>It&#8217;s good to have a heuristic feel for the danger\/flexibility\/noise sensitivity (synonyms) of each statistical learning tool. I roughly have this spectrum in my head:<\/div>\n<div>\n<\/div>\n<div style=\"text-align: center;\"><em>Very specific, structured, safe<\/em><\/div>\n<div style=\"text-align: center;\"><\/div>\n<blockquote>\n<div style=\"text-align: center;\">Optimize 1 parameter, require crossvalidation<\/div>\n<div style=\"text-align: center;\">\u2193\n<\/div>\n<div style=\"text-align: center;\">Optimize 2 parameters, require crossvalidation<\/div>\n<div style=\"text-align: center;\">\u2193<\/div>\n<div style=\"text-align: center;\">Optimize parameters with too little data, require regularization<\/div>\n<div>\n<div style=\"text-align: center;\">\u2193<\/div>\n<\/div>\n<div style=\"text-align: center;\">Extrapolation<\/div>\n<div>\n<div style=\"text-align: center;\">\u2193<\/div>\n<\/div>\n<div style=\"text-align: center;\">Nonlinear (SVM, tree bagging, etc)<\/div>\n<div>\n<div style=\"text-align: center;\">\u2193<\/div>\n<\/div>\n<div style=\"text-align: center;\">Higher-order variable dependencies<\/div>\n<div>\n<div style=\"text-align: center;\">\u2193<\/div>\n<\/div>\n<div style=\"text-align: center;\">Variable selection<\/div>\n<div>\n<div style=\"text-align: center;\">\u2193<\/div>\n<\/div>\n<div style=\"text-align: center;\">Structure learning<\/div>\n<\/blockquote>\n<div style=\"text-align: center;\"><\/div>\n<div style=\"text-align: center;\"><em>Very general, dangerous in noise, voodoo<\/em><\/div>\n<div style=\"text-align: center;\">\n<\/div>\n<div>This diagram is worth expanding. If anyone has any suggestions, please leave them. <\/div>\n<\/div>\n<div class=\"blogger-post-footer\"><img loading=\"lazy\" loading=\"lazy\" decoding=\"async\" src=\"https:\/\/blogger.googleusercontent.com\/tracker\/3965329713014965566-8331717359424403673?l=www.maxdama.com\" alt=\"\" height=\"1\" width=\"1\"><\/div>\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>I used to be very gung-ho about machine learning approaches to trading but I&#8217;m less so now. You have to understand that that there is a spectrum of alpha sources, from very specific structured arbitrage opportunities -&gt; to stat arb -&gt; to just voodoo nonsense. As history goes on, hedge funds and other large players [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_seopress_robots_primary_cat":"","_seopress_titles_title":"","_seopress_titles_desc":"","_seopress_robots_index":"","footnotes":""},"categories":[2,4],"tags":[714,356,77],"class_list":{"0":"post-2532","1":"post","2":"type-post","3":"status-publish","4":"format-standard","6":"category-business-intelligence","7":"category-data-mining","8":"tag-data-sets","9":"tag-machine-learning","10":"tag-modeling"},"amp_enabled":true,"_links":{"self":[{"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/posts\/2532"}],"collection":[{"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/comments?post=2532"}],"version-history":[{"count":0,"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/posts\/2532\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/media?parent=2532"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/categories?post=2532"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.smartdatacollective.com\/wp-json\/wp\/v2\/tags?post=2532"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}