{"id":17265,"date":"2023-08-09T16:31:00","date_gmt":"2023-08-09T20:31:00","guid":{"rendered":"https:\/\/demo2.dedicatedhost247.com\/crypto\/how-to-build-your-own-bitcoin-language-model\/"},"modified":"2023-08-09T16:31:00","modified_gmt":"2023-08-09T20:31:00","slug":"how-to-build-your-own-bitcoin-language-model","status":"publish","type":"post","link":"https:\/\/demo2.dedicatedhost247.com\/crypto\/how-to-build-your-own-bitcoin-language-model\/","title":{"rendered":"How To Build Your Own Bitcoin Language Model"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<p><em>This is an opinion editorial by Aleksandar Svetski, author of \u201cThe UnCommunist Manifesto\u201d and founder of the Bitcoin-focused language model Spirit of Satoshi.<\/em><\/p>\n<p>Language models are all the rage, and many people are just taking foundation models (most often ChatGPT or something similar) and then connecting them to a vector database so that when people ask their \u201cmodel\u201d a question, it responds to the answer with context from this vector database.<\/p>\n<p>What is a <a href=\"https:\/\/learn.microsoft.com\/en-us\/semantic-kernel\/memories\/vector-db\">vector database<\/a>? I\u2019ll explain that in more detail in a future essay, but a simple way to understand it is as a collection of information stored as chunks of data, that a language model can query and use to produce better responses. Imagine \u201cThe Bitcoin Standard,\u201d split into paragraphs, and stored in this vector database. You ask this new \u201cmodel\u201d a question about the history of money. The underlying model will actually query the database, select the most relevant piece of context (some paragraph from \u201cThe Bitcoin Standard\u201d) and then feed it into the prompt of the underlying model (in many cases, ChatGPT). The model should then respond with a more <em>relevant<\/em> answer. This is cool, and works OK in some cases, but doesn\u2019t solve the underlying issues of mainstream noise and bias that the underlying models are subject to during their training.<\/p>\n<p>This is what we\u2019re trying to do at Spirit of Satoshi. We have built a model like what\u2019s described above about six months ago, which you can go try out <a href=\"https:\/\/www.app.spiritofsatoshi.ai\">here<\/a>. You\u2019ll notice it\u2019s not bad with some answers but it cannot hold a conversation, and it performs really poorly when it comes to shitcoinery and things that a real Bitcoiner would know. <\/p>\n<p>This is why we\u2019ve changed our approach and are building a full language model from scratch. In this essay, I will talk a little bit about that, to give you an idea of what it entails.<\/p>\n<h2>A More \u2018Based\u2019 Bitcoin Language Model<\/h2>\n<p>The mission to build a more \u201cbased\u201d language model continues. It\u2019s proven to be more involved than even I had thought, not from a <em>\u201ctechnically complicated\u201d<\/em> standpoint, but more from a <em>\u201cdamn this is tedious\u201d<\/em> standpoint.<\/p>\n<p><em>It\u2019s all about data.<\/em> And not the quantity of data, but the quality and format of data. You\u2019ve probably heard nerds talk about this, and you don\u2019t really appreciate it until you actually begin feeding the stuff to a model, and you get a result\u2026 which wasn\u2019t necessarily what you wanted.<\/p>\n<p>The data pipeline is where all the work is. You have to <em>collect<\/em> and <em>curate<\/em> the data, then you have to <em>extract<\/em> it. Then you have to programmatically <em>clean<\/em> it (it\u2019s impossible to do a first-run clean manually). <\/p>\n<p>Then you take this programmatically-cleaned, raw data and you have to <em>transform<\/em> it into multiple data <em>formats<\/em> (think of question-and-answer pairs, or semantically-coherent chunks and paragraphs). This you also need to do programmatically, if you\u2019re dealing with loads of data \u2014 which is the case for a language model. Funny enough, other language models are actually good for this task! You use language models to build new language models.<\/p>\n<figure>\n        <img loading=\"lazy\" src=\"https:\/\/bitcoinmagazine.com\/.image\/c_fit%2Ccs_srgb%2Cfl_progressive%2Cq_auto:good%2Cw_620\/MTk5OTM5MTgwOTI5ODIwMjg4\/language-models-all-the-way-down.jpg\" height=\"338\" width=\"620\"><\/p>\n<\/figure>\n<p><em>Then<\/em>, because there will likely be loads of junk left in there, and irrelevant garbage generated by whatever language model you used to programmatically transform the data, you need to do a more intense <em>clean<\/em>.<\/p>\n<p><em>This<\/em> is where you need to get human help, because at this stage, it seems humans are still the only creatures on the planet with the agency necessary to differentiate and determine <em>quality<\/em>. Algorithms can kind of do this, but not so well with language just yet \u2014 especially in more nuanced, comparative contexts \u2014 which is where Bitcoin squarely sits.<\/p>\n<p>In any case, doing this at scale is incredibly hard unless you have an army of people to help you. That army of people can be mercenaries paid for by someone, like OpenAI which <a href=\"https:\/\/www.crunchbase.com\/organization\/openai\">has more money than God<\/a>, or they can be missionaries, which is what the Bitcoin community generally is (we\u2019re very lucky and grateful for this at Spirit of Satoshi). Individuals go through data items and one by one select whether to keep, discard or modify the data.<\/p>\n<p>Once the data goes through this process, you end up with something clean on the other end. Of course, there are more intricacies involved here. For example, you need to ensure that bad actors who are trying to botch your clean-up process are weeded out, or their inputs are discarded. You can do that in a series of ways, and everyone does it a bit differently. You can screen people on the way in, you can build some sort of internal clean-up consensus model so that thresholds need to be met for data items to be kept or discarded, etc. At Spirit of Satoshi, we\u2019re doing a blend of both, and I guess we shall see how effective it is in the coming months.<\/p>\n<p>Now\u2026 once you\u2019ve got this beautiful clean data out the end of this \u201c<em>pipeline,<\/em>\u201d you then need to <em>format<\/em> it once more in preparation for \u201c<em>training<\/em>\u201d a model.<\/p>\n<p>This final stage is where the graphical processing units (GPUs) come into play, and is really what most people think about when they hear about building language models. All the other stuff that I covered is generally ignored.<\/p>\n<p>This home-stretch stage involves training a series of models, and playing with the parameters, the data blends, the quantum of data, the model types, etc. This can quickly get expensive, so you best have some damn good data and you\u2019re better off starting with smaller models and building your way up.<\/p>\n<p>It\u2019s all experimental, and what you get out the other end is\u2026 <em>a result\u2026<\/em><\/p>\n<p>It\u2019s incredible the things we humans conjure up. Anyway\u2026<\/p>\n<p>At Spirit of Satoshi, our result is still in the making, and we are working on it in a couple of ways:<\/p>\n<ol>\n<li>We ask volunteers to help us collect and curate the most relevant data for the model. We\u2019re doing that at <a href=\"https:\/\/repository.spiritofsatoshi.ai\/\">The Nakamoto Repository.<\/a> This is a repository of every book, essay, article, blog, YouTube video and podcast about and related to Bitcoin, and peripherals like the works of Friedrich Nietzsche, Oswald Spengler, Jordan Peterson, Hans-Hermann Hoppe, Murray Rothbard, Carl Jung, the Bible, etc.\n<p>You can search for anything there and access the URL, text file or PDF. If a volunteer can\u2019t find something, or feel it needs to be included, they can \u201cadd\u201d a record. If they add junk though, it won\u2019t be accepted. Ideally, volunteers will submit the data as a .txt file along with a link.<\/p>\n<\/li>\n<li>Community members can also <a href=\"https:\/\/www.spiritofsatoshi.ai\/#\/help-training\">actually help us clean the data, and earn sats<\/a>. Remember that missionary stage I mentioned? Well this is it. We\u2019re rolling out a whole toolbox as part of this, and participants will be able to play \u201cFUD buster\u201d and \u201crank replies\u201d and all sorts of other things. For now, it\u2019s like a Tinder-esque keep\/discard\/comment experience on data interface to clean up what\u2019s in the pipeline.\n<p>This is a way for people who have spent years learning about and understanding Bitcoin to transform that \u201cwork\u201d into sats. No, they\u2019re not going to get rich, but they can help contribute toward something they might deem a worthy project, and earn something along the way.<\/li>\n<\/ol>\n<h2>Probability Programs, Not AI<\/h2>\n<p>In a few previous essays, I\u2019ve argued that \u201cartificial intelligence\u201d is a flawed term, because while it <em>is<\/em> artificial, it\u2019s <em>not<\/em> intelligent \u2014 and furthermore, the fear porn surrounding artificial general intelligence (AGI) has been completely unfounded because there is literally no risk of this thing becoming spontaneously sentient and killing us all. A few months on and I am even more convinced of this.<\/p>\n<p>I think back to John Carter\u2019s excellent article <a href=\"https:\/\/barsoom.substack.com\/p\/im-already-bored-with-generative\">\u201cI\u2019m Already Bored With Generative AI\u201d<\/a> and he was so spot on.<\/p>\n<p>There\u2019s really nothing magical, or intelligent for that matter, about any of this AI stuff. The more we play with it, the more time we spend actually building our own, the more we realize there\u2019s no sentience here. There\u2019s no actual thinking or reasoning happening. <em>There is no agency<\/em>. These are just \u201cprobability programs.\u201d<\/p>\n<p>The way they are labeled, and the terms thrown around, whether it\u2019s \u201cAI\u201d or \u201cmachine <em>learning<\/em>\u201d or \u201cagents,\u201d is actually where most of the fear, uncertainty and doubt lies. <\/p>\n<p>These labels are just an attempt to describe a set of processes, that are really unlike anything that a human does. The problem with language is that we immediately begin to anthropomorphize it in order to make sense of it. And in the process of doing that, it is the audience or the listener who breathes life into Frankenstein\u2019s monster.<\/p>\n<p><strong>AI has <em>no<\/em> life other than what you give it with your own imagination.<\/strong> This is much the same with any other imaginary, eschatological threat.<\/p>\n<p><em>(Insert examples around climate change, aliens or whatever else is going on on Twitter\/X.)<\/em><\/p>\n<p>This is, of course, very useful for globo-homo bureaucrats who want to use any such tool\/program\/machine for their own purposes. They\u2019ve been spinning stories and narratives since before they could walk, and this is just the latest one to spin. And because most people are lemmings and will believe whatever someone who sounds a few IQ points smarter than them has to say, they will use that to their advantage.<\/p>\n<p>I remember talking about regulation coming down the pipeline. I noticed that last week or the week before, there are now \u201cofficial guidelines\u201d or something of the sort for generative AI \u2014 courtesy of our bureaucratic overlords. What this means, nobody really knows. It\u2019s masked in the same nonsensical language that all of their other regulations are. The net result being, once again, \u201cWe write the rules, we get to use the tools the way we want, you must use it the way we tell you, or else.\u201d<\/p>\n<p>The most ridiculous part is that a bunch of people cheered about this, thinking that they\u2019re somehow safer from the imaginary monster that never was. In fact, they\u2019ll probably credit these agencies with \u201csaving us from AGI\u201d because it never materialized.<\/p>\n<p>It reminds me of this:<\/p>\n<figure>\n        <img loading=\"lazy\" src=\"https:\/\/bitcoinmagazine.com\/.image\/c_fit%2Ccs_srgb%2Cfl_progressive%2Cq_auto:good%2Cw_620\/MTk5OTM5MTgwOTI5NzU0NzUy\/awake-yet.jpg\" height=\"777\" width=\"620\"><\/p>\n<\/figure>\n<p>When I posted the above picture on Twitter, the amount of idiots who responded with genuine belief that the avoidance of these catastrophes was a result of increased bureaucratic intervention told me all that I needed to know about the level of collective intelligence on that platform.<\/p>\n<p>Nevertheless, here we are. Once again. Same story, new characters.<\/p>\n<p>Alas \u2014 there\u2019s really little we can do about that, other than to focus on our own stuff. We\u2019ll continue to do what we set out to do.<\/p>\n<p>I\u2019ve become less excited about \u201cGenAI\u201d in general, and I get the sense that a lot of the hype is wearing off as people\u2019s attention moves onto aliens and politics again. I\u2019m also less convinced that there is something substantially transformative here \u2014 at least to the degree that I thought six months ago. Perhaps I\u2019ll be proven wrong. I do think these tools have latent, untapped potential, but it\u2019s just that: latent. <\/p>\n<p>I think we have to be more realistic about what they are <em>(instead of artificial intelligence, it\u2019s better to call them \u201cprobability programs\u201d)<\/em> and that might actually mean we spend less time and energy on pipe dreams and focus more on building useful applications. In that sense, I do remain curious and cautiously optimistic that something does materialize, and believe that somewhere in the nexus of Bitcoin, probability programs and protocols such as Nostr, something very useful will emerge. <\/p>\n<p>I am hopeful that we can take part in that, and I\u2019d love for you also to take part in it if you\u2019re interested. To that end, I shall leave you all to your day, and hope this was a useful 10-minute insight into what it takes to build a language model.<\/p>\n<p><em>This is a guest post by Aleksander Svetski. Opinions expressed are entirely their own and do not necessarily reflect those of BTC Inc or Bitcoin Magazine.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/bitcoinmagazine.com\/culture\/how-to-build-your-own-bitcoin-language-model\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>This is an opinion editorial by Aleksandar Svetski, author of \u201cThe UnCommunist Manifesto\u201d and founder of the Bitcoin-focused language model Spirit of Satoshi. Language models are all the rage, and many people are just taking foundation models (most often ChatGPT or something similar) and then connecting them to a vector database so that when people ask their \u201cmodel\u201d a question, it responds to the answer with context from this vector database. What is a vector database? I\u2019ll explain that in more detail in a future essay, but a simple way&hellip;<\/p>\n","protected":false},"author":1,"featured_media":17266,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[30],"tags":[45,619,6689,1868],"_links":{"self":[{"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/posts\/17265"}],"collection":[{"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/comments?post=17265"}],"version-history":[{"count":0,"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/posts\/17265\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/media\/17266"}],"wp:attachment":[{"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/media?parent=17265"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/categories?post=17265"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/demo2.dedicatedhost247.com\/crypto\/wp-json\/wp\/v2\/tags?post=17265"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}