Home / Transcripts / Simulations Plus, Inc. (SLP) · September 30, 2020

Simulations Plus, Inc. (SLP) Earnings Call Transcript

September 30, 2020

NASDAQ US Health Care Health Care Technology special 67 min

Earnings Call Speaker Segments

Arlene Padron executive
#1

Good morning, and thank you for joining us. There is no debate here. You are in the right place to learn more about AI-driven drug design. A few housekeeping notes. For optimal video and audio connection, we recommend you close any additional web or streaming application that may affect your Internet bandwidth. We take your privacy right seriously. And by attending this event or participating in the upcoming Q&A session, you are allowing us to contact you for follow-up. This webinar is being recorded for future playback on our website and YouTube channel. So if you need to drop at any time, you can come back and watch on demand. [Operator Instructions] We'll strive to have at least 10 minutes of Q&A after the presentation and live software demonstration. Now before I hand over the control to Eric Jamois, I have a quick question for you. Are you currently using ADMET Predictor software? We'll give you just a few seconds here. [Voting]

Arlene Padron executive
#2

I'm going to go ahead and close the poll now. All right, it looks like a few of you are new to ADMET Predictor software. So it's my pleasure to introduce Eric Jamois, Business Development Director. Eric, take it away.

Eric Jamois executive
#3

All right. Thank you, Arlene. Good morning, good afternoon, good evening, everyone. On behalf of Simulations Plus, I welcome you to today's webinar, where we will present and demonstrate our new AI-driven drug design technology. As you've seen in one of the slides, it's termed AIDD and is delivered under our new ADMET Predictor 10.0 platform that was released just a couple of weeks ago. We have a lot of information to share. So without further ado, I'm introducing Dr. Marvin Waldman, Senior Research Fellow at Simulations Plus. Marvin has been on the R&D team of Simulations Plus for about 14 years and he's responsible for most of the research and development work that has gone into this new AIDD technology. And Michael Lawless, Senior Principal Scientist, will also be joining us as a panelist for this event. So Marvin, the floor is yours.

Marvin Waldman executive
#4

Thank you very much, Eric. Good morning, good afternoon or evening, everyone, depending on where you're located. So I'm going to be talking today, introducing our new AI drug discovery module with ADMET Predictor 10, or APX as we like to call it. Before I get started, we're going to provide a bit of motivation and background for how we got interested in this and also some background on prior work in this area by others. So a few years ago, we initiated an in-house project to develop novel antimalarial concept -- compound with a proof of concept to show how our in silico models and other capabilities could be used to design compounds for drug discovery. This has recently been published in the Journal of Molecular -- Computer-Aided Molecular Design, appeared a couple of months ago. The work was headed up by Bob Clark, another Research Fellow -- Senior Research Fellow at Simulations Plus, and Michael Lawless, who's a panelist, and a number of collaborators. The basic idea was we took compounds with known antimalarial activity against one of the target receptors, PfDHODH, and built activity models based on those compounds, and you can see the performance of those models here. We generated analogs from some of those known compounds using our in-house SMIRKS transforms capabilities and then tweaked those analogs a bit to make them a bit more novel and/or synthetically tractable. And also based on their predicted properties, we decided -- [ concentrated out ] to have some of those compounds synthesized. And then we had those properties measured such as solubility, logP, pKa. We measured the properties of those compounds and then compared them to our predictions. And generally, the results were very favorable. We then actually had those compounds evaluated against the parasite, a Plasmodium parasite, in terms of its ability to inhibit the growth of that parasite. And some of the compounds that were most active ex vivo were also the ones that were predicted to have the highest inhibition against the PfDHODH target. So that was really encouraging. We also actually included pharmacokinetic simulations to make sure the compounds that -- were showing good pharmacokinetic properties as well. And the last sentence on this article, I encourage you all to take a look at it, was the following: It seems likely that further iterations and in vivo characterization would be productive. Further iteration. So this led us to begin thinking about can we automate this process? You'd really like to do this over and over again, and you can't do the experiments in -- experimental evaluations iteratively on a computer, but you can certainly do the analog generation and predictions iteratively. Of course, this idea of designing compounds on the computers, it's not new to us. It's a -- from my perspective, it's been around since at least the early 1990s. Some of the first such programs that did this were basically trying to dock fragments from a fragment library into a target receptor, such as LUDI and MCSS/HOOK and SPROUT. Later on, I'd say it became more recognized widely within the pharmaceutical community that "it ain't just activity anymore." This became popularized with Lipinski's Rule of 5 and the recognition that ADMET was a key aspect of designing drugs. And I think a wider recognition of the drug design process is really a multi-objective problem. In the early to late 2000s, programs started to appear that took this into account, multi-objective ligand and structure-based design. Often, the multi-objective aspects were combined into a single function, say, by adding up different objectives or a weighted average or a geometric mean. And you can find examples of this in programs such as EA-Inventor, Muse, and there are a number of publications where pharma companies developed in-house programs along these lines. And in the last decade, I think more widely Pareto-based optimization has come into play rather than simply combining all the objectives into a single function. The recognition of perhaps Pareto-based optimization, which I'll be explaining more in a minute, is an alternative or perhaps better route. And also even in the last few years when we first started this project, there was maybe one publication that I was aware of from the [ culture group ], but lately there have been, I would say over the last few years, hundreds of publications on using deep learning and generative algorithms to either generate SMILES streams or generate molecular graphs with deep learning neural networks. I'm not going to be talking much more about that. We are aware of that. It's obviously an area of active research at the moment. But at least for the time being, we've chosen to use built-in technology we already had in place based on SMIRKS transformations to do the compound generation. So how do you generate virtual molecules on a computer? Various ways have been tried. For example, you can start with a given molecule and make what I would call elementary changes by changing a single atom or a single bond. Or you could add common fragments to the molecule, take a hydrogen and substitute it with a carboxylic acid. Or take a carboxylic acid and substitute it with a tetrazole, for example. And then when you're doing these kinds of things, you need to be aware of synthetic feasibility issues. It's very easy, as I'll be showing in a few minutes, for the computer to generate compounds that are totally nonsensical. So some of the things that have been tried were to -- had been to add filters based on structural alerts. You don't want the computer generating peroxides, presumably hemiacetals. Lots of other things that are problematic. Another way of doing this has been using a library of known synthetic reactions and building blocks. Typically, this was based on SMIRKS transformation language, which I'll give you some examples of. They're developed by Daylight for representing chemical reactions. And they're too limited here with just known synthetic reactions. You may have some limitations in terms of chemical diversity and novelty. And finally, again, most recently, generative algorithms based on deep learning generating SMILES or graph representations of molecules. Secondly, we have the issue of how to -- okay, you generate molecules but now you want to evaluate them so that you generate good molecules. So as I said, they're now -- it's recognized that this is a multi-objective problem. You could use a weighted sum or some other kind of combining function. Pareto optimization has become also a popular way of doing this. And the question then is what criteria to use as your multi-objective criteria. So typically, you'd want to have some models for representing activity, perhaps docking into receptor and scoring that. You'd want to take account of ADMET liability. Sometimes, similarity to a known lead has been used. So you'd want to design compounds that are similar but not exactly the same as some known active. You need to take account of synthetic accessibility. The compounds need to be drug-like, and you need to worry about the chemical filters. Okay. So with that as the background, I'll introduce how we actually chose to do this within our AIDD program. So here's the general workflow. You start with one or more seed molecules that are used to initial -- as initial starting structures. You then select from these molecules at random a molecule which you then randomly select from a list of chemical transforms represented as SMIRKS-based, and you generate candidate molecules. You then evaluate their properties, for example, ADMET Risks, synthetic difficulty, monomer activities, pharmacokinetics using our HTPK module. And then from these molecules, you prune out the list, you prune them down to molecules which are good. And by good, I mean they sit on what's called the Pareto front. So in this small graphic, from a set of molecules in this 2-dimensional space, the ones you select, assuming you wanted to minimize both properties and drive compounds from the lower left corner, would be the ones in this lower left region of this graph, okay. Once you've done that pruning, you generate more analogs. You can set up chemical transforms and basically regrow the population with candidates, evaluate their properties, prune it again, wash and repeat N times, and it's usually a fairly large number. And gradually, what you find when you do this is that your population molecules gradually grows towards this lower left hand region of this graph toward desirable properties. Okay. So for generating analogs, we've developed a library of what we would call chemically intelligent SMIRKS transforms. And why do I say intelligent? So here's an example of the SMIRKS transforms. Suppose you wanted to represent something where I'm going to take something which is not a fluorine, an element, to transform and turn it into a fluorine. In SMIRKS language, this would be -- could be represented in this -- with this syntax: #9 stands for atomic number 9, fluorine, and exclamation point means not atomic number 9, take an atom and then change it to something which is atomic number 9, fluorine. That's the very simple way of telling the program to take something which is not a fluorine and make it a fluorine. The problems arise because that it's easy to make compounds which are undesirable. It's kind of a transform. For example, if I took a carboxylic acid and converted the OH group to a fluorine, I would produce an acid halide, which is almost certainly something you don't want to have in a drug compound and probably hardly anything that you would expose to a human. So we improve our SMIRKS pattern in this case, for example, by saying, okay, it's not a fluorine. It's a degree 1 compound. It's attached to a carbon. We don't want to produce OF either, for example. But it's not -- notice the bank sign, and this represents a -- what's called a [ carrier 6 marks ] but it's not something which is attached to a carbon and doubly -- which is itself doubly bonded to O, N or S, so avoids acid halides, [ thiol ] acid halides and so on. And if it satisfies all of these criteria, okay, then go ahead and convert it to a fluorine. And currently, we have about 150 transforms like this all with various criteria to avoid -- as best we can, avoid producing compounds that are undesirable while still performing these kinds of transforms. So here's a sample of that list. I'm not expanding everything that we categorize in. At the user level, you can turn on or off individual transforms. You can turn on and off whole categories. We have very elementary things, as I was showing, like CF3 to methyl or add CF3 or add amide or add a bromo, for example, and more complicated things like aromatize a 5- or 6-membered ring or change a pyridone to a phenyl and vice versa and so on. So we have both very simple things and some of the more complex things, but all things that we feel are reasonable medicinal chemistry transforms to apply to a molecule. Okay. So -- and now onto the properties. This is the first screenshot you see when you bring up the AIDD workflow. It allows you to select what properties to optimize, what objectives to use in the optimization. In this example, selected are synthetic difficulty scoring function, fraction bioavailable from using our HTPK capabilities, ADMET Risk, and one of our built-in models for inhibition -- or inhibiting HIV-1 strand transfer. You can select whether you want a given property to be minimized or maximized. You can select whether to perform pharmacokinetic simulations for rat or human, what dose level to use. And there's various other things you can modify by editing a file that contains the parameters for performing HTPK. And there's also penalties you can apply, which I'll speak to in a few minutes, involving status quo factors. I'll explain that in a few moment. Okay. Now I was mentioning -- okay. So there's -- coming back to this [ for a moment ]. There's 50 models here that you can select from. These are not all of our ADMET models. We don't include classification models. So these are basically a set of regression models. But our recommendation actually is for most of our models to make use of the ADMET Risk score rather than the individual models. The reason for this is severalfold, but the main one is if you use too many objectives, that actually creates a problem when you're using Pareto optimization. So our typical recommended use for this protocol would be something like 4 to 5 objectives for a [ sustainable ] ADMET Risk, perhaps synthetic difficulty, 1 or 2 activity models, for example, one to model activity, perhaps one to model [ seal ] activity, and probably good pharmacokinetics, such as bioavailability. And now why not too many objectives? So here's an example. I showed in that small screen, showed a little -- it expanded now. There's 1,000 points that could represent molecules in a 2-dimensional space. Two objectives in this case represented from randomly distributed points -- random and normally distributed points along each direction. And if you apply Pareto selection to these points as a -- you would get 7 points in our Pareto optimal color here in amber. Okay. If you -- in experiments we've done, if you expand this to 5 dimensions, just 5 dimensions, now out of 1,000 points, 100 are Pareto optimal because as you expand the dimensionality, you find more and more points tend to sit on the surface. More and more points will at least have one parameter that's very good and that are not dominated. What do I mean by Pareto optimal here? Pareto optimal -- these points in the lower left-hand corner are Pareto optimal. That means that there is no point among the remaining points that is better than in both properties. For a point to be better than, say, one of these amber points, it would have to sit -- would have to be placed to both the left and below this point so that it would be better in both properties. As you can see here, there are no points that dominate these points that are better than in both properties while these points dominate all the remaining points. So with -- for any given blue point here, you'll find one of these amber points is better than them -- at least one of the amber points is better than them in both properties. And so these are -- so the so-called Pareto optimal points. But too many objectives leads to choosing too many molecules, and that eventually becomes undesirable, as I said, because most of them are good in just 1 or 2 properties. So what we recommend is to make use of our ADMET Risk scoring, which basically combines aspects of absorption, distribution, toxicity and metabolism on the various criteria into a set of rules, motivated and similar in some ways to Lipinski's Rule but making use in most cases of our ADMET property models. And depending on your point of view, either violating or passing this rule or triggering this rule causes a penalty rate to apply -- to be applied to the score. So the higher the score, just like Lipinski's Rule of 5, the worst the molecule is deemed to be from an ADMET perspective. The second thing, the second objective that we think is unique to us, at the moment at least, is the ability to use PBPK, a group of PBPK simulations based on our state-of-the-art, leading GastroPlus software, which makes use of the ACAT model. It's a compartmental model to perform pharmacokinetic simulations. And it's very fast to run. As we've mentioned earlier, run, generate -- run these simulations on hundreds of molecules per second. And the other unique thing about it is it's purely in silico, as implemented in ADMET Predictor. All of these properties needed, permeability, solubility, pKa, logP, logD, are generated using our in silico model. So no experimental quantities are needed, which makes it suitable for use of the AIDD protocol because we're generating virtual molecules. We don't know what their properties are yet. No experimental properties are required. It's extremely rapid and we made it even more rapid by implementing it using multi-threading. So it can run in parallel, essentially. [ That's giving you ] the only calculations themselves now. As to synthetic accessibility and difficulty, that's based on this paper from Peter Ertl and Schuffenhauer that was published in 2009. It's based on generating -- taking a molecule and representing it as circular fragments starting from each -- from a given central atom and going up to 3 bonds out and taking those fragments and comparing them to the frequencies that the fragments appear in a sample subset of the PubChem database. Ertl used about 1 million compounds. And the more frequent the fragment is found to appear in this PubChem database, the more it is assumed that, that is a fragment that's readily synthesizable. So it's deemed to be more easy to synthesize. On top of that, you add some various complexity terms, involve a number of atoms, number of macrocycles, stereocenters, spiro centers, bridging atoms and so on. And he came up with the scoring function, which ranges in his implementation from 1 to 10. We did something fairly similar, except we expanded that out and sampled actually 47 million compounds, virtually the entire PubChem after you throw out things containing metals and BaP structures, valence violations and things like that. One other thing we did was in Ertl's implementation, the atoms on the outer layer of the circular fragment are represented as fully wild-carded. And we chose to make a slight distinction there in terms of whether they're aromatic or aliphatic, thinking that can make a difference to really how you would go about synthesizing it, how synthesizable it is. Complexity is the same, and we chose to make the range 0 to 10. And here, we're showing a comparison of Ertl's score on the horizontal axis versus our function on the vertical axis based on 40 compounds Ertl published and then the subsequent publication that also gave the scores for 18 -- [ maybe ] 12 more compounds, a total of 52 compounds -- 50, 54 compounds. And you can see the scores are fairly similar. Then we also -- this is something you'll see coming into play a little later. If we evaluate the range of those scores, a subset of the WDI -- and the subsetting is performing an analysis very similarly the way Lipinski generated a subset of the WDI for generating the Rule of 5. So on these 2,260 drug-like subset of the WDI, the scores peak around 2.5, and somewhere above 4 to 5 is about as far as you want to go before it's clear that these are things that you probably -- [ the left ] high scores are probably things you don't want to think about designing for a drug especially in terms of the effort it might take to make a compound that you don't already know is active based upon predictive models. So we don't want to go much about it more generally. Probably you would find in medicinal chemistry are carbon-resistant [ 10 C ] compounds that score much above 4. But then we found when we actually evaluated this on virtual compounds that the synthetic difficulty turned out to be a bit too optimistic on some virtual molecules. So we decided to augment it with drug-like filters based on some publications. Think of chemical functional group, chemical moieties that would be undesirable in drugs, done in a way so that the maximum penalty we applied is 4 even if it's got a lot of violations. Because once, as I said, you go much above 4, it doesn't really matter. You probably don't want that compound anyway. And this augmented version, we found a -- the AIDD process. And we used the augmented version instead of just adding another Pareto objective because we're trying to somewhat limit the number of objectives for reasons I mentioned earlier. So here are some early results that illustrates some issues we found at the beginning. If you look at some of these molecules, I'm pretty sure you would agree that these are not things you want to show to your medicinal chemist colleague that's something you should consider making. This is what came out of some of the early runs where we were trying to optimize HIV inhibition and evaluating it for synthetic difficulty and ADMET Risk. And you can see that from the algorithm's perspective, this is a very nice-looking compound. It's got phenomenal predicted HIV inhibition, fairly low ADMET Risk. And from the synthetic difficulty score, it doesn't look all that bad, but I would say to any chemist's eyes, it looks pretty horrendous. So when we took a closer look at this, we discovered, not too our -- much to our surprise, that for all these -- most of our models -- well, the ones you'll be able to slack from, we opine that we have something that evaluates it from -- in terms of its applicability domain. And all of these compounds were found to be outside the applicability domain of the HIV inhibition model. So what is applicability domain? So the way we evaluate that is for the model -- the compounds that are used in the training set, we evaluate what ranges of the descriptors are used to train the model, the property model, the activity model, and we basically draw a box here in 2 dimensions, in multi-dimension, the hyper box. And then we expand it to allow for a little bit of buffer room by 10%. And so anything that's sitting inside this buffer zone, this dashed square box, is said to be in scope. And if the properties are sitting outside that box, it's out of scope. It's outside the applicability domain. And so what we came up with was we're going to add a penalty for compounds that are out of scope. And the default is 10. And so we subtract 10 from the score. Now these compounds don't look so good anymore. We also added a similar penalty to ADMET Risk. So we call -- temporarily called it ADMET Risk plus. And so -- but this -- since ADMET Risk consists of a set of rules, the penalty is done on a per rule basis. And as I mentioned, we also added penalties for synthetic difficulty for things that have undesirable fragments. And when you add -- put all of those together, now these compounds don't look so good, which is somewhat reassuring. There's a kind of a flip side to that coin in terms of synthetic difficulty. Compounds can also be very easy to make that have synthetic difficulty scores of 0 or very low, yes. Well, they're very easy to make. They have kind of pretty low ADMET Risk, not surprisingly. They're not that great from an HIV perspective, but because of the goodness of these 2 properties, like one that's sitting on the Pareto front, they're kind of hard to eliminate. So we came up with the idea of assigning capping values. The idea is if -- so here, I'm showing it for synthetic difficulty in fraction bioavailable. But if the compound is better than 2.5 in terms of its synthetic difficulty -- remember, that was the peak on that WDI-focused subset, well, we say that's good enough. We don't care if it's better than that, and we want you to focus on the other objectives, okay, and similarly for fraction bioavailable in this case. If it's greater than 90% bioavailable, that's good enough. We don't want to keep compounds that have some really low synthetic difficulty because they're better on that score but bad on the other objectives. So the way we deal with that is by assigning this capping value. These are optional, but we have some recommended -- recommendations here. And when we apply that strategy to these very simple compounds, so all their synthetic difficulty score suddenly jumped to 2.5. Now when you compare them to these other reasonable compounds, somewhat more reasonable at least, I would say, sort of drug-like looking, that also came out of the algorithm. They're assigned synthetic difficulty scores of 2.5. But now that's better once you sort of discard that synthetic difficulty by capping it in terms of their predicted inhibition, and their ADMET Risk is actually better. And that allows these compounds to dominate these from a Pareto perspective. They're better on 2 scores. They're tied on the other, and so they're considered better and then they cause these compounds, thankfully, to get thrown out of the population. They don't pollute your population. You don't waste time generating analogs or compounds you're not interested in, and it basically improves the overall performance of the algorithm. Okay. So the next screenshot you would see allows you to do further restrictions on the chemical space. You can specify scaffold queries. You can specify that I want compounds that satisfy a certain chemical scaffold. You can also specify the bulk filtering criteria that becomes hard filters, not just penalty filters. We've provided it with a default file for this. You can either use it, clear it out, not use it, modify it, use your own. The filters can be either, for example, SMARTS pattern. So in this example, we're saying this carbon, carbon double bond O to nitrogen, is basically an amide functionality. We say we don't -- this yield grade, we're able to [ form ] if I don't want more than 3 amide groups in my molecule. So you're gradually able to form -- match this query -- matches this query. That means that's a problem, basically. So you're going to allow up to 3 amide groups in the molecule. Or I say this example is not a SMARTS pattern. These are prebuilt in properties you can use. So I don't want more than 65 atoms in my molecule -- 65 heavy atoms. The scaffold query can either be drawn using our MedChem Designer drawing tool. And you can also just have that wild card or create list of atoms or bond types. So at these positions, I'm saying I'll only allow a carbon or a nitrogen in this phenyl ring. I'm going to let a carbon or oxygen at this substitution point. And this bond could be single or double. Or you can do something much more complicated using SMARTS, which I don't have time to go into, but it's described -- the SMARTS syntax is described in our user manual and also on Daylight's website. And finally, the last dialog specifies run parameters. Most of them are, I think, clear enough from earlier parts of the presentation, how many generations they want to run, number of candidate molecules to generate per generation, which would be the size of the initial population. Since the transforms are generated randomly, you can specify at random seed. 777 happens to be my favorite random seed. And sometimes, if you're generating thousands of generations, these -- the runs can take hours to days potentially. So we allow you to specify an intermediate file frequency when you can -- it will print out the results of that run. It will write a SMILES file with the results of the population after so many generations, and you can examine those results while the calculation is still running to see if it's producing something desirable. If you want to stop it at some point, maybe need to tweak the parameters a bit, maybe still producing undesirable molecules in some aspect, you want to change the filters and the scaffold and so on. And for reproducibility, there's a parameter specified, whether you want it running multi-threaded. So in terms of -- the one parameter I didn't mention was its minimum size. So as I mentioned earlier, you can have up to 7 -- in this example, you can only get 7 molecules out of 1,000, but suppose you wanted at least 500. So the way we implement that is, as I mentioned earlier, if you don't satisfy this minimum size criteria, then after it selects the first set of Pareto optimal and the minimum size is not specified, it will go back and select from the remaining compounds what's left as Pareto optimal. We call that Pareto layers. So you're sort of like peeling the surface of an onion. You peel like layer after layer from the Pareto surface from the candidate molecules and reinforce this only after 50% the generations had been run because we don't want to start off generating -- putting in lots of bad molecules into the population. But after about 50%, most of the molecules are pretty good at that point. And as you start peeling off from the Pareto surface, there, you'll tend to get a good molecule. So in summary, what we've developed is a highly automated and its fully customizable protocol, which is now commercially available in APX, making use of about 150 chemically intelligent transforms, which are further customizable and user-controllable. You can add. You can turn off our transforms. You can add your own transforms. Property objectives based on up to 50 built-in property models that are available, including ADMET Risk, which automatically combines them for you, includes high throughput HTPK simulations, fully automated, fully in silico, the ability of user models that you can build with ADMET Modeler to represent activity or other properties. Many options for controlling/limiting the chemical space to avoid chemical nonsense coming out, such as a scaffold definition, chemical filters, the various out-of-scope penalties, the augmented synthetic difficulty and performance, which I didn't speak to until now, but running it on my laptop, which is an i7 8-core laptop for physical or virtual cores running 7 to 8 threads, we can generate and evaluate up to about 10 million molecules per 24 hours. So that would represent, using default settings in the program, about 2,000 generations per day. We're current -- this is currently being evaluated and validated. We just put out a press release about this and the collaboration with a large pharmaceutical company where they supplied us compounds to build -- a data set to build activity models, which we then used with the rest of the algorithm to design novel compounds, which would provide candidate novel compounds that the algorithm generates that they have been synthesizing and testing. And this is an ongoing collaboration and iteration to see what we can come up with. And finally, I'm going to now provide a working example, a demo. You can see that this is real and produces some fairly interesting results. We're going to start with BACE1 inhibitors starting from a known active seed. BACE1 is a receptor -- enzyme receptor that has been used and is being investigated as a potential target for treating Alzheimer's disease. So the basic idea was we took 370 known BACE inhibitors from literature, and we built an activity model using ADMET Modeler. We added an additional risk-type penalty or model, let's say, based on representing risks associated with blood-brain barrier penetration since we're targeting Alzheimer's here. We added scaffold and additional filtering criteria based on the pharmacophore presented in the literature for BACE1. And in the example, we're going to illustrate how the algorithm is able to perform scaffold hopping to find alternative scaffolds and, in fact, find compounds that have been previously synthesized and tested and scaffolds that have been made. So the motivation or inspiration for this came from an article published a few years ago by a group of workers at Janssen. Rombouts is the first author. And their goal was they started with this compound that was a known inhibitor but was not BBB penetrant mainly because it was too basic. And the goal was to reduce the basicity of the compounds and resulted in a compound that was BBB penetrant. But it turns out the activity they got, 7.6 nanomolars, is possible to get better activity. And some of the compounds from the literature data set had better activity, but they have other issues. We're going to work from this pharmacophore. This is a representation of pharmacophore they showed in an article that we're going to create a scaffold that represents the analogs used to generate this [ related series ], starting from a phenyl ring which allows some substitution here, connected to an amide group, connected to another phenyl ring, connected to another ring containing an amidine group, which makes interactions with the BACE1 receptor here in this amide nitrogen, also makes interactions here. And we're going to allow for some limited substitution along these rings but limited so that they don't start to clash with the pockets of the receptor. We've added an additional risk model that represent BBB, making use of our in-house BBB filter model and logBBB model. So if these predict that it's problematic in terms of BBB penetration, that causes this rule to be triggered. Also, if it's a Pgp substrate, that would cause another rule to be triggered or violated or -- and finally, to deal with the basicity issue, too basic a pKa, we deal with that in terms of a descriptor we have for representing what fraction of the molecules are positively charged cationic at physiological pH. And that ranges from 0.5, representing a monoprotic pKa of about 7.4 to 0.975, which would be a monoprotic pKa of 9. So this -- above a pKa of 9 or fraction cationic above that 0.975, that would fully trigger this rule. And below 0.5, you would get no penalty at all. We built an activity model from these 370 compounds, and I guess now I'm ready for the demo. Okay. So these are the 370 compounds that we built -- or that were used to define the model. And now I'm going to do a query. I've already preloaded a query because that represents the scaffold in terms of its SMARTS pattern. It also restricts it in terms of limiting how much substitution you can do around those phenyl rings that I mentioned earlier. And if I search for that, we find at least 370 molecules from the literature that we extracted, 17 of them are satisfying that scaffold query. And so I'm going to hide everything other than those ones. And if we go down -- and these are already sorted by their measured pIC50 against the BACE1 target. And this is the compound from the Rombouts paper that they synthesized. This is the most active compound synthesized in that paper. We actually see that in the literature, there's another compound that satisfies that scaffold that is actually more active but it has potential BBB issues. It's predicted to be a Pgp substrate and its fraction cationic is a bit high. So what we're going to do is we're going to start with this compound as the seed for running AIDD, and I'm going to launch the AIDD wizard. I've already preloaded it with all the settings. I want here, so we're including synthetic difficulties, fraction bioavailable, ADMET Risk, the BACE1 model, productivity and the BBB risk. And I've also set some capping values here. As I mentioned earlier, 2.5 is a reasonable value to use, 90 for fraction bioavailable. If it's below 1, let's say 0.9 for ADMET Risk, we'll consider that good enough. And in terms of activity, 10 represents 0.1 nanomolar and we'll be fine if it's 0.1 nanomolar or above. So these represent values like if you're better than these, we want you to focus on the other objectives. We have already preloaded the scaffold query. This is the input file containing filter criteria. It's slightly modified from the default one we have because some of the rules involving [ rings themselves ] need to be modified to accommodate the number of rings that are already in the scaffold. And so -- and finally we have -- we come to the run setting. I've also turned off a number of transforms here. I'll show you that. All the transforms that make rings were turned off because that would create groups that are clashing with the pockets of the receptor. And we turned off adding rings here. We turned off some other things. So these are all things that -- they would be filtered out anyway, in all likelihood, but it's just to prevent the thing from generating groups that when you substitute around these rings would cause clashes with the receptor. So we don't bother making them. But we're still allowing a lot of transforms here, change its chain length, right, or even add functional groups, still quite a few things that are turned on, okay. So still using my favorite random number seed. We've reduced the number of optimization generations that runs in a reasonable time, number of candidate molecule size -- minimum size to 200, and I'm going to launch it. And it's launching. This will take about 10 minutes to run. So we're not going to wait for it to complete. I'm going to show you what it looks like after it does complete, but I'll let it run for a few generations so you'll get an idea. It's running single thread for product generation because we need to do that to make the results reproducible. In general, for a real production run, you probably don't care about that. Now these are generating ADMET and HTPK properties, running multi-threaded and it's running pretty fast and then the synthetic difficulty. And then those have already pruned the Pareto population and now it's generating more products. So far it's got 1,000 molecules left. It's Pareto pruning and generating 12 molecules after the first generation. Now it's on generation 2, generating more products. And after generation 3, it jumped to 46 molecules, of which 39 are new. And I'll let it run for one more generation. And now it's up to 96 molecules, okay. So while that's running, I'm actually going to load the results. So these are what you get after it finishes. It takes about 10 minutes to run. These are the resulting molecules. It found 423 molecules altogether after 50 generations, actually 422 and we added back in the seed molecule. You can see how it compares. Now I'm going -- so -- and we show the results of the objective. So in terms of predicted inhibition, values range from 4.7 to almost 10; ADMET Risk, 0.9, that was --- because it was capped there up to almost 7; BBB risk, 0; a full value of 3 in synthetic difficulty. Now I'm going to filter that to remove things we don't really care about. So we want at least 8 for inhibition. We want no more than 1.9 for ADMET Risk. We want really low BBB risk. And we don't want to go above 4 for synthetic difficulty, okay. And the result of that produced 36 molecules after all that filtering. So if I focus on that by hiding everything that's unselected -- let's do that again. So I accidentally selected something. So actually, we already have 36 molecules, I apologize, [ I didn't do much I think ]. Okay. So if we sort then by activity and we go down here, we've seen some interesting compound, sulfones. In fact, let's take a look at that and let's launch a query and let's look for sulfones, O equals S equals O. And 18 compounds out of the 36 have ringed sulfone group in them. If we go back to that original data set that was used in the training -- sorry, I need to filter this or this scaffold query. We see that, in fact, this was the compound here that was used as a seed. The compound that was more active was a sulfone. In fact, there are only 2 sulfones in this entire data set. One of them has satisfied the scaffold and one doesn't. If I take this compound and copy it and go back to the results -- and let's turn off the filters. So we set the ranges back to 0. Nothing is hitting now. But if I'm going to launch a query, it takes [ one click more ]. This is the query, that molecule that was in the original training set, the algorithm found it. The algorithm actually was able to jump from the seed molecule back to that compound, the most active compound in the training set that satisfied that scaffold query. It was filtered out earlier, as already mentioned, because it had this high BBB risk. But it found other compounds that had this kind of a ring system, the sulfone in the ring. And so then it found 18 compounds that satisfied that scaffold and that satisfied the filtering criteria I was applying that are analogs of this. The compound jumped from the seed compound here into this scaffold. It might manage to modify the ring and [ bounce ] alternative compound. We've also found 18 more compounds, 1 of which are in the training set that have a single -- that have this ring scaffold and different analog substitutions around this ring. All of which it found were actually more desirable because they didn't have high BBB risk. So if we go back and look at our sulfones again here, we actually have found 124. But if I filter them out, as I did earlier, there are things that are actually desirable. And we start down -- we actually [ plan out ] the compounds with quite high activity, good bioavailability, low ADMET Risk, low BBB risk and synthetically tractable. That actually did better than any of the compounds in the training set or in the compounds that were synthesized by the Rombouts group. And finally, one other interesting query I would like to show is the [ following ] group. So let's go back here, turn on the filters again. Every -- there's nothing -- everything is shown. We launch one more query. This is another kind of scaffold. Let's focus on that. There's another scaffold, an interesting one, where amide group attached to the amidine group, also fairly active and meeting the other criteria, low BBB risk and so on. These scaffolds don't appear in the training set. If I go back to the training set data, we find everything. Oops, sorry. Okay. Perform the same search for this group, I don't find anything. It's on the scaffold that was not in the training set. In fact, the eighth compound in this list right there with scaffold fluorochlorine. It's in fact this compound that was all earlier synthesized by Janssen. So again, it jumped from this oxygen-type compound to another scaffold, which is known to be an active -- was an -- also an active compound found and it found in total of at least an [ an inhibited 10 ]. Here's the [ control table ] filter ranges, search for the scaffold again. It found 13 compounds that match this scaffold, and some of them have actually better predicted activity than the [ one that's remained ]. And with that, I will conclude the presentation and the demo. Hopefully, this has been somewhat convincing about the power of this program. And I guess we'll now open it up for questions.

Arlene Padron executive
#5

Thank you, Marv. Before we get started with the questions, we do have another question for the audience. Do you currently have internal efforts or interest directed towards generative chemistry for compound optimization? While we continue to get your questions and to the question panel and Michael gets ready to begin the Q&A session, you can submit your answer now. Just a few more seconds here. [Voting]

Marvin Waldman executive
#6

The current running -- it's still running but it's almost done.

Arlene Padron executive
#7

All right. So it looks like we do have quite a few people with internal efforts. So with that, I will now turn it over to Michael.

Michael Lawless executive
#8

Thank you, Arlene. Okay. So Marvin, the first question that's queued up here is the following. There seems to be an assumption that the properties used in the multi-objective optimization are independent from one another. But something like fraction bioavailable normally depends upon the dose, which in turn depends at least on activity and probably on solubility, too, unless a very basic, fully dosed proportional PK model is assumed, which from what is being described now seems not to be the case. Any comment on that aspect?

Marvin Waldman executive
#9

We enable you to choose the dose level in [ the trial ] if I launch the fraction bioavailable.

Michael Lawless executive
#10

Percent add. Okay. Thank you so much.

Marvin Waldman executive
#11

So you can choose the dose level. You can also control other parameters of the simulation. In a normal run, you can actually choose multiple dose levels for AIDD to choose the dose level, but we set it at 10, which is a fairly large dosing level. There is, of course, some correlation between things like solubility, logP and pharmacokinetic simulation. The point of the pharmacokinetic simulations is they take account of the way these properties combine in such a way that it wants them to be desirable in ways that the various combinations of these properties produce good pharmacokinetics. So it's the kind of the whole point of the pharmacokinetics. It isn't just the properties in isolation but the way they combine to give the compound its good pharmacokinetics. And so that's why we provided those additionally.

Michael Lawless executive
#12

Okay.

Marvin Waldman executive
#13

Or you could -- okay. That's my answer for [ the trial ]. I'm sticking with it. Okay. By the way, the other run finally finished. And just so you can see, it produced the same 423 molecules. The results are what I was showing earlier.

Michael Lawless executive
#14

Cool. Okay. And this is just kind of a clarification. Is the extremely fast, no data required PBPK model the same as the HTPK model or something different?

Marvin Waldman executive
#15

It's exactly the same.

Michael Lawless executive
#16

Okay. When -- okay. So another question, when combining scores, have you considered using consensus methods?

Marvin Waldman executive
#17

That's something you can do. We haven't tried it but we have, let's call -- you can specify what we call an MLR or PLS model where you can use various ways of combining different functions or descriptors. And most recently, we've implemented something of a maximum in -- so you can take the maximum or the minimum of a set of values, which is sort of like a consensus. Or you can take a geometric mean or an arithmetic mean. So we -- I mean ADMET Risk is kind of a consensus because it's taking an arithmetic sum of various criteria. And yes, it's -- things along that line are possible by the user by making use of the tools we already have. But beyond ADMET Risk, we haven't tried a lot of examples along -- I guess along those lines. But there are certainly -- the program -- the customizability of the program allows me to do things like that.

Michael Lawless executive
#18

Okay. Thank you. So another one is, can you comment on the pros and cons of SMIRKS-based transformations versus the deep learning-based generative model?

Marvin Waldman executive
#19

I would say the main pro is that you have more control over what kinds of things it's going to generate versus the deep learning ones admittedly are very powerful but it's kind -- from everything I've seen to date, it's -- they're kind of hard to control in terms of what kinds of molecules they're going to generate. So that's I would say the advantage of the SMIRKS transforms. Its potential disadvantage is you have to specify them all out to generate that level of control. You have to make sure you've got what you want there, and the deep learning ones are kind of figure it out for themselves in a way. But as far as I can tell, there's still a lot of active research in terms of playing around. First, it was all based on the SMILES streams. Now more recently, there's generative graph versions of them. And so the main disadvantage, I haven't seen this offered yet in a commercially available tool. Some companies are offering it through consulting services and things like that, but I'm not aware of it being available as a general commercial tool. And the second issue is having control over what action produces versus being -- it's quite a black box at the moment from my perspective.

Michael Lawless executive
#20

That's a good answer. Okay. So another question. It seems arbitrary decisions were taken, property caps, 10% outside model space regarded as in-scope. Have these been validated?

Marvin Waldman executive
#21

The 10 -- I mean to some extent, a lot of science involves making your risk-based decisions, I would say. I don't know how you exactly would validate the 10%. It's somewhat of a choice. With regard to the specific capping values and so on I used in the program, we did -- if -- I did play around with changing them a bit. And if you make modifications to them, if you could change the random number seed and so on, you won't get exactly the set of molecules. You won't get -- you may or may not get exactly the compounds that were synthesized. But what I did find is you do get those scaffolds, the scaffolds of -- that represent the 2 forms -- the 2 classes of the compounds that were synthesized show up. They may show up a little later in the population and they show up a little earlier but they do show up. As you get those specific compounds or not -- they might get Pareto optimized out by other compounds ultimately. So they may or may not show up, but the scaffolds, trying to -- plays around with those parameters tweaking into other values that still seem to show up. So there's a certain level of stability to the algorithm with a certain level of -- it will vary. I would say the fact that it varies from run to run in some ways but you don't punt the results always to be the same. You want to see how much things can change. And so I mean that's how much we've done so far. And we've got another example. The longer you run it, the more -- let me have one more thing. The longer you run it, if you run it for thousands of generations, eventually, the results tend to converge somewhat to similar answers. But of course, if you're playing around with the capping values and things like that, you will get somewhat different answers, but that, I would say, is a strength of the algorithm. You can play around, try to tweak these things and see how much they affect the results and if they improve the results. Okay. That's my answer.

Michael Lawless executive
#22

Okay. That's a good answer. So just a question on the activity model. So could you give a little bit more information on how the activity models are developed and which type of models they are? So maybe just talk about your COX-2 model -- excuse me, the base...

Marvin Waldman executive
#23

They're -- well, anyway, they're artificial neural network ensembles. So they make use of our ADMET Modeler program. So if I were to take this data set, let's then hide everything and I were to launch it here on the spreadsheet rows -- cancel, sorry about that. This is the -- what's interesting, oh, because of a different descriptor. It actually does. Anyway, I think I need to move the seed compound. Well, anyway, let me just load a different data set so you're going to get an idea, [ a different 10 or 3D ]. Hang on real quick. So if we calculate descriptors first, it says we're able to thread it. So there are 104 molecules that's generated in a couple of seconds. Then we can launch our ADMET Modeler program until we get ADMET Modeler [ with lock key ]. And this is how -- so it's done for the 370 but [ I ran it ] very fast. So I'm trying to build a model with logP based on 104 measured values. I'm going to do some artificial neural networks. I'm going to let it automatically choose sets of descriptors and sets of numbers of neurons to build a grid of the model and will quickly build -- this took a little longer with 370 and I played around various things. And you can look at model for some of these and so on. So this is building models with different neural network architectures, in this case based on 104 molecules of measured logP values, and you can build different artificial neural network ensemble models. And that's how that was done. The test selection is automatic, and that's how that was done but that's 370. Anyway, we don't have -- time doesn't permit me going through the details of everything we did there, but that's the general idea. We use our ADMET Modeler program to build an artificial neural network ensemble model for the baseline activity.

Michael Lawless executive
#24

Perfect, perfect. Okay. And then another question is, are all the molecules found at the end of the Pareto -- or end of the process Pareto optimal? Which they are, I'm sure.

Marvin Waldman executive
#25

Yes, they are.

Michael Lawless executive
#26

Doesn't that force -- okay. Doesn't that force the user to explore a large number of possibilities? Is there no other way to meaningfully rank them?

Marvin Waldman executive
#27

Well, not -- yes, many of them are Pareto optimal but often, they're Pareto optimal in just 1 or 2 objectives. So the point I was trying to show earlier, I go back to the results, and the actual results was to use this filtering because yes, we have 423 molecules here, but many of them are just good in one property. So if we filter them out here, I'd say I want it to be at least 10 nanomolar active, I don't want things with bad ADMET Risk, okay? This was the whole point. I don't want things that have bad BBB risk, okay? I don't want things that are hard to make, okay? So it will produce things that may be bad in several objectives but good -- just good in one is the nature of the Pareto optimization. But then when you apply this filtering criteria at the end, now out of these 423 molecules, I only have 36 and I can study these and decide which one of these look interesting potentially to make, test, whatever and so on. So that's the approach we would recommend, yes, to get many at the end in a Pareto optimal. Rather than ranking them, we would suggest to filter them so that they have all -- they have, let's say, acceptable properties across the entire spectrum of objectives, okay. That's my answer.

Michael Lawless executive
#28

Okay. I think that's a good answer. And then the last question I have. Is this module already included in ADMET Predictor?

Marvin Waldman executive
#29

Well, it's included in the latest release. It's an optional purchase in APX or ADMET Predictor 10. It was just released about a couple of weeks ago. So it is available as of now, yes.

Michael Lawless executive
#30

Perfect. All right. That's all the questions we have. Let's see. Do I turn it over back to Arlene or -- okay.

Arlene Padron executive
#31

Yes. Thank you. Thank you, Michael, Marv, and thank you, Eric. You've definitely shown us how the AIDD module integrates ADMET Predictor's top-ranked ADMET property prediction models with multi-objective compound optimization capability. Now for our audience to test drive ADMET Predictor's next-generation technology for yourself, please visit our website at www.simulations-plus.com. This concludes our webinar for today, and we look forward to virtually seeing you next time at the next online event. Have a great day.

Read the full transcript via the API

You're viewing the first half of this call. Get the complete Simulations Plus, Inc. transcript - plus 252,000+ transcripts from 12,000+ companies, speaker segments and full-text search - through the EarningsAPI REST API or hosted MCP server.

Get an API key View API docs →

For developers and AI pipelines

Programmatic access to Simulations Plus, Inc. earnings transcripts and 252,000+ others is available through the EarningsAPI REST API and the hosted MCP server. Quarterly plans from $105 - full transcripts, speaker segments, full-text search, and the /api/v1/transcripts/recent polling endpoint for ETL pipelines.