Let's cut through the jargon, myths, and nebulous world of data, machine learning and AI. Each week we'll be unpacking topics related to the world of data and AI with the awarding-winning team at 1000ML. Whether you're in the data world already or looking to learn more about it, this podcast is for you.
Don't forget to hop on over to 1000ml.io and ApogeeSuite.com, as well as our FB and LinkedIn communities.
In the ever-evolving world of technology, ethics play a pivotal role in shaping the future. Recently, an intriguing development caught our attention – Canada's AI Code of Conduct. In this article, we'll dissect the key points discussed in our latest podcast episode, shedding light on the significance of these regulations for the field of artificial intelligence.
In the ever-evolving landscape of technology, artificial intelligence (AI) stands as a beacon of promise, holding the potential to transform industries and redefine our daily lives. In this article, we delve into two thought-provoking pieces that shed light on distinct aspects of the AI realm. The first discusses the challenges startups face in realizing the AI dream, while the second provides an early view of the future of generative AI through compelling charts.
In the ever-evolving landscape of technology, businesses are navigating uncharted territories, seeking innovation and efficiency. A recent McKinsey & Company article sheds light on the transformative potential of Generative AI, emphasizing its role as a catalyst for innovation in the corporate realm. This technological marvel goes beyond automation, providing businesses with a creative force capable of generating ideas, designs, and even human-like text.
In the fast-paced world of technology, one sector that has often been viewed as traditional and resistant to change is the legal industry. However, in recent years, this sector has witnessed a transformation powered by Natural Language Processing (NLP), a branch of artificial intelligence focused on enabling computers to understand and generate human language. This innovation is revolutionizing the legal landscape, streamlining processes, and opening up new possibilities for the legal profession.
Natural Language Processing (NLP) is not just a buzzword but a revolutionary technology. This article serves as a beginner's guide, delving into the world of NLP, its significance, and its transformative potential.
Generative AI has captured the imagination of many, promising to be a technological marvel that can mimic human creativity, generating text, art, and even music. But beneath the surface of this digital wizardry lies a shadow—a pressing ethical conundrum. In this article, we'll delve into the depths of Generative AI and explore the potential dangers, misuse, and the critical importance of ethical guidelines.
In the realm of artificial intelligence, Generative AI has emerged as a transformative force, pushing the boundaries of what machines can create. Over the past decade, this field has witnessed remarkable advancements, thanks to the collective efforts of AI researchers, engineers, and data scientists. To demystify the intricate world of Generative AI and explore its boundless potential, we turn to the insights and perspectives of leading experts in the field.
Generative AI has emerged as a groundbreaking force in the realm of technology, enabling machines to craft content, replicate human cognitive processes, and even produce awe-inspiring works of art. As we find ourselves on the cusp of a new era, the future of generative AI promises nothing short of spectacular. This episode delves into the visions of prominent researchers to provide a glimpse of where Generative AI is heading and the ethical considerations that must accompany its progress.
Welcome to "Tech on Trial," your ultimate source for unraveling the dynamic interplay between technology and the legal realm. Join us as we explore the revolutionary impact of Artificial Intelligence (AI) and Natural Language Processing (NLP) on legal workflows. In this exciting episode, we're diving headfirst into a transformation that's reshaping the way legal professionals work, paving the way for enhanced efficiency, accuracy, and strategic decision-making.
In today's digital age, where technology is reshaping industries at an unprecedented pace, the legal profession stands at the threshold of a transformational journey. At the heart of this evolution lies the realm of Natural Language Processing (NLP), an innovative branch of artificial intelligence that bridges the gap between human language and computer understanding. In this episode, we embark on an exploratory voyage into the world of NLP, uncovering its profound implications for legal professionals and the intricate tapestry of the law.
In the dynamic landscape of technology, the legal sector is experiencing a groundbreaking transformation thanks to AI-powered solutions. In this episode of Tech on Trial, we delve into the revolution taking place in the realm of legal document search and retrieval. Joined by Victor, the visionary CEO of Apogee Suite, we explore how AI is reshaping the way legal professionals navigate the complex world of documents.
In a world where technology continually reshapes industries, the legal field is no exception. Welcome to the riveting realm of "Tech on Trial," a podcast that delves into the convergence of law and technology, unraveling the innovations that are propelling the legal profession into uncharted territories.
Welcome to "Tech on Trial," the podcast where we explore the dynamic intersection of technology and the legal world! I'm your host, Sal, and today's episode promises to be a fascinating deep dive into the realm of legal tech. In this blog article, we'll uncover the transformative impact of Artificial Intelligence (AI) and Natural Language Processing (NLP) on document creation in the legal industry.
In this fast-paced digital age, the legal landscape demands innovation and efficiency, and that's where NLU comes in. NLU is like a wizard of language processing, enabling machines to comprehend human language nuances contextually, going beyond simple word matching. It's a game-changer for legal research, e-discovery, contract analysis, predictive analytics, and beyond!
The legal landscape is undergoing an incredible metamorphosis, catalyzed by the integration of Artificial Intelligence (AI) and Natural Language Processing (NLP) technologies. These cutting-edge innovations are fundamentally changing how legal professionals conduct research, assess contracts, and draft legal documents.
In the fast-paced world of legal practice, managing vast volumes of documents efficiently can be a daunting challenge. Lawyers and legal professionals often find themselves buried under piles of paperwork, searching for the proverbial needle in a haystack. However, a groundbreaking solution has emerged - AI-driven document classification - poised to transform the legal landscape.
In our latest episode of "Tech on Trial," we explored the fascinating world of Natural Language Processing (NLP) and its revolutionary impact on e-discovery processes for law firms. NLP is not just a buzzword; it's a game-changer that empowers legal professionals to navigate vast volumes of electronic data with unprecedented efficiency and accuracy.
Welcome to the "Tech on Trial" blog, where we explore the latest tech trends and innovations that are reshaping the world around us. In this episode, we delve into the incredible potential of AI-powered natural language generation (NLG) in document creation.
Welcome to Tech on Trial,the podcast where we put cutting-edge technologies to the test. In this episode, your host, Sal, is joined by the brilliant CEO of Apogee Suite, Victor, to explore the fascinating world of natural language understanding (NLU) in contract management systems. The podcast aims to delve into the latest advancements in the tech industry, and today, they're focused on how NLU is revolutionizing the way businesses handle their contracts.
Welcome to another episode of Tech On Trial!
Today we will discover how AI and NLU technologies enhance legal research and analysis, enabling smarter decision-making.
Stay tunned!
Welcome to Tech On Trial! Your new podcast focused on how Technology can be used in the Legal Industry.
In this first episode, we are covering how you can use NLP in the Legal Industry, especially when it comes to document analysis!
In today's episode of All Things Data, we delve into the world of contract management and how it can significantly impact businesses. Contract management, with the help of automation, artificial intelligence (AI), and natural language processing (NLP), streamlines workflows and enhances productivity. By optimizing contract processes, businesses can achieve cost savings, and revenue generation, and minimize legal liability and risk. In this blog article, we explore the key benefits of contract management automation and highlight the role of Zenith, an innovative contract management suite.
Today, we focus on the role of AI and Natural Language Processing (NLP) in fraud detection. Fraud detection involves building a cyclical feedback loop to continuously improve detection methods over time. When it comes to content that can be analyzed using NLP, the first step is understanding the text. By mining the data and developing a fraud dictionary based on patterns and keywords, potential signals for fraud can be identified. This initial understanding is then used to create models and evaluate their performance.
Today, we'll delve into the realm of legal and judicial AI and explore the possibilities of developing an AI program capable of predicting legal decisions and case outcomes. This field presents a ripe opportunity for automation, as it involves a common and repeatable process that legal professionals undertake when handling cases.
In the realm of artificial intelligence, Natural Language Understanding (NLU) has emerged as a crucial component in various applications, transforming the way we interact with machines. One fascinating area where NLU is making significant strides is in Document AI. By enabling machines to comprehend and extract meaningful insights from unstructured text documents, NLU in Document AI is revolutionizing information processing, data analysis, and decision-making. In this episode, we'll explore the transformative potential of NLU in Document AI and delve into its practical applications.
Discover how AI and NLP technologies are revolutionizing supply chain management in our latest blog article, "Streamlining Supply Chain Processes: The Role of AI and NLP." Dive into the world of artificial intelligence and natural language processing and explore their pivotal role in optimizing supply chain operations.
In the realm of legal artificial intelligence (AI), understanding and extracting relevant clauses from legal documents are essential tasks. However, the real challenge lies in comprehending the meaning and context of these clauses.
In the realm of Natural Language Understanding (NLU) and Natural Language Processing (NLP), contextual analysis plays a vital role in deciphering the meaning and intention behind a piece of text. Without context, words and sentences can be subject to multiple interpretations, leading to ambiguity. In this article, we will delve into the significance of context in NLP and explore how it is analyzed to enhance text understanding.
Legal documents can be complicated and time-consuming to process, but with the advancement of Natural Language Processing (NLP) and Artificial Intelligence (AI), companies have been able to work with them more efficiently. In this episode, we will delve into how legal document processing works usingNLP and AI.
We'll explore how NLP can revolutionize businesses by improving employee efficiency and enhancing the commercial side. Specifically, we'll focus on the main uses of NLP in large businesses.
RPA, AI, and NLP are three technologies that can work together to automate repetitive tasks and create a more efficient system. In this podcast, we will provide an overview of these technologies and discuss how they can be combined to benefit organizations.
Document AI is a product suite offered by 1000ML that enables businesses to extract valuable insights from their documents using natural languageprocessing (NLP) techniques. Unlike traditional document processing systems, Document AI employs semantics and context to understand the desired output,rather than being limited by pre-existing rules.
Today we will be discussing the lifecycle of AI projects and how they are undertaken, both internally and externally. The focus will be on why certain internal AI projects fail to meet the expected outcomes, and how organizations can avoid falling into these traps.
One of the biggest factors affecting the success of internal AI projects is the quality of data. Gathering complete data, including metadata, is crucial todeveloping accurate AI models.
In today’s episode, we will talk about NLP and its use in business from an executive's point of view. NLP is explained as the process of extracting data from language to be used by computers and AI programs to create decisions or predictions, and the process involves extracting content from documents, understanding the meaning of the content, analyzing the content, and making a decision.
By integrating AI and NLP with ERP, organizations can achieve significant benefits such as reducing manual processes, improving accuracy, making better decisions, increasing efficiency, productivity, and cost savings, identifying new growth opportunities, minimizing risk, and improving their bottom line. As unstructured data grows, AI and NLP will become increasingly important for organizations that want to remain competitive.
Extracting clauses from legal documents is a complex task that requires an understanding of the legal domain and its technical language. Machine learning and natural language processing (NLP) programs are used to understand the meaning of clauses and classify them. The most commonly used language model in the legal domain is the FILAC model, but it has limitations and does not help with crafting arguments or understanding more about the case. To overcome these limitations, the team at 1000ML developed a better classification system.
Advancements in Natural Language Processing (NLP) and Artificial Intelligence (AI) have made it possible for companies to automate the process of analyzing legal documents, which are often lengthy and complex. The process involves document ingestion, where the algorithms can extract entities, understand the structure and formatting of the document, and summarize the document to provide an overview of its contents and classify it into topics such as lawsuits or rental disputes.
Today we will talk about AI in Legal Organizations when it comes to classification and partitioning clauses. By organizing your documents you can interact with them in a way that you have a full knowledge of the context and information that a document has.
Today we will talk about AI, NLP & RPA interests with each other. There are several advantages of combining all these technologies into a system in order to create the perfect AI environment and pipeline.
Today we will talk about the RPA can optimize processes in legal operations. In the legal world, being a document-search based there is room for optimization and automation, and here is where RPA enters. These systems help you reduce cost and time, without substituting other systems you might have.
Join us as we explore the power of ChatGPT and its impact on document analysis in the Apogee Suite. Our expert guests will delve into the different models of ChatGPT and how they can enhance creativity and language understanding. Discover how ChatGPT can help you ask the right questions for your document analysis process and why context is so critical in this field. Tune in to our podcast for a deep dive into the world of ChatGPT and document analysis.
Welcome to our latest episode of the podcast where we'll be delving into the world of NLU in Document AI.
We'll start by discussing Optical Character Recognition (OCR) and how it differs from Natural Language Understanding (NLU). OCR allows for the transformation of non-digital formats into digital ones, but it does not provide understanding of the content. NLU, on the other hand, goes beyond OCR by providing a readable version for the computer, allowing it to understand any kind of document.
Traditionally, understanding documents was a rule-based process that relied on the word order in a phrase. However, this system lacked semantic understanding and was not portable, requiring the building of new systems for each industry, resulting in longer and less accurate processes.
At 1000ml, we've worked with a variety of industries to adapt Contract AI, a field of Document AI, to their specific needs. For example, in the pharmaceutical industry, we built a pipeline for categorizing invoices and contracts according to the medicine they bought. In the legal industry, we worked on a system that allowed NLU to focus on patent documentation. In the financial and insurance industries, our work centered around form information extraction, resulting in faster and more accurate processes. And in healthcare, we combined different information about a patient in one place.
All paper-driven industries can benefit from Document AI, especially when NLU is involved. Join us next week as we continue to explore the exciting world of NLU in Document AI.
We started the year talking about NLP, and to continue it’s mandatory to talk about context, the contextual analysis of a text. Contextual analysis is understanding where a sentence is coming from.
When we think about how computers understand human language, we have to understand that it's a process and not an easy one also each language has a sense and syntax model that defines each one. So in order to process and understand that information, computers need signals to guide the process and to associate the language to action themselves.
We're at the tail end of the year now, and we're just talking about what 2023 may hold for 1000ml and generally in the world of NLP.
So 1000ml has a suite of NLP and AI products that we called Apogee. The suite is comprised of just general text ingestion that can be documents, which is usually what people go for but also allows you to ingest any website or webpage basically whatever you want that has text in it, including.
The crux of it is that Apogee suite builds up a series of pipelines and APIs including a recommendation engine and relevancy engine API that allows all the tools that we build on top of it to utilize these engines to properly search for content. So you might think of things like, okay, well, I go online and I do a Google search and I'm looking for a flight or something.
Imagine all kinds of descriptions, text messages, and emails that you send internally, giving you the ability to ask content questions. So imagine a use case like you're in e-commerce and you're looking for the right type of pants. Somebody's like, I need black pants, you can go all the way from like athleisure to like super serious, like tuxedo pants if you're a man. So are you just gonna surface everything along the lines and how do you know exactly the kind of things that they're looking for?
ADA sort of allows you to ask your database, your text, anything. The curated data searching that we've done prior with just our semantic search is now getting a big boost, allowing us the ability to really ask the document things, so not just metadata things like who's the author, which basically any system could do.
When will 1000ml tackle as ambitious Witness Prep AI, come out of the world of us doing a lot of work in legal and judicial and deeply understanding the paradigms that exist there with the advent of us creating semantic search and then layering on ADAs so you can ask the content anything and you could read any content. So again, Witness Prep AI is all about ingesting a lot of things about a specific case or a specific area of law or kind of context that may occur.
It's largely going to be an exercise for the enterprise world until we really understand how to maximize its value and use and then can reduce it to a point where it's then manageable for people who want it as an add-on into our main suite, whether that's for our regular B2C clients or for clients and customers.
So that's a bit of a view into 2023 and what 1000ml is gonna be working on. It's still all about NLP and AI and mostly focused on Apogee Suite. It's our big framework and it's what we sell the most of.
2022 has been a fun year in AI and NLP, and today we thought we'd take a minute and reflect back on all the things that have occurred. Some of the important things that have occurred in the world of NLP and a lot of the generative world of NLP in 2022.
GPT has been a big deal in NLP, Generative, Pre-trained Transformer is a generation technology to help us generate usually text content and that can really be like all kinds of things. If you go look on YouTube and you kind of like dabble around with GPT-3 OpenAI, you'll see examples of all kinds all the way from like, give me a tweet basically, you know, like generate a tweet from you, which is kind of interest.
So for a machine to generate and do it well and in within the context of what you want is actually really cool and really good. So this year, obviously there's been advances in GPT-3 they did allow for the editing and insertion of data, like human data back into the reports. So if you were to ask GPT to specifically generate something for you, could actually embed new information in there to make it slightly better is a big deal.
Another thing that they've done since they noticed in the past that was happening is they've added a lot of plagiarism detection into their generation. They've done that with slightly better word parsing and quite a bit of synonym generation so that the content will be as unique as possible and avoid plagiarism.
Another really big thing done this year was to allow for the monitoring or ongoing review of data repositories that allow GPT to constantly scan and add new content as it becomes available and then make GPT's engine more relevant to your search, as well.
Tune in next week we're gonna talk a little bit about our roadmap for 2023 and what we see could happen in NLP next year.
Happy 2023 everyone!
We started to wrap up the year a little bit talking about the considerations that you have to think of when you're deciding whether to build or buy AI. That really feeds in well and quite nicely into today's topic where we're gonna talk more about the technical knowledge required for AI projects. So this is assuming you went through the framework of deciding whether to build or buy, and you were like, you know what? We can build this, let's do this ourselves.
So now that you're in the world of let's build this because you've decided that we have the staffing resources, and we're able to get consultants who could do this, having enough know-how and enough people with the breadth of knowledge required for this. Also, we've worked with or have the infrastructure available for this.
Generally, the organization is in a good shape, or our change management processes are quite good, understanding how to make sure that this blends into our organization well having in mind the total cost of ownership of such a project.
For a very long time anybody who's done a lot of work in NLP has chosen to usually start their journey with the package NLTK is largely like the go-to and it is the basis, it's the foundation that provides quite a bit of functionality, but some of the things that are critical for you to know if you're going to do serious NLP you need to know stemming.
You need to split the content either into phrases because you may want to analyze phrases or sentences, or split it into paragraphs, but usually into words in order to understand and know how to use a part of speech triggers bringing you endless opportunities.
For example, there's the possibility of doing a strictly extractive summary where you're going to pull things out of a document. Imagine an entire thesis, there is an abstract, and using an NLP on that whole thesis, you'd probably just pull that abstract becoming your summary.
.So you want to make an abstract, so you're largely looking to take the knowledge and understand it and then abstract the knowledge of that document. So you're really summarizing as opposed to pulling specific parts of the content out.
Again, high risk for it cause we're really good at this stuff. But you could also use machine learning and AI as a sort of input to NLP. You'd find that there are many clusters when you're doing this that can then help lead you to create an AI program based on those clusters.
Lastly, if you wanted to use the NLP as an input to different programs, you could do that for things like the things we do, which we do a lot of work in document intelligence, in contract AI.
Today we're gonna talk about whether to buy or build AI systems.
As opposed to the shorter write-ups that we usually have, you'll probably get a lot of benefit from rating it and even using it as notes when you are making the decision on whether to build or buy a new system generally, not just AI.
Like building or buying most technology systems or new systems for our company you have to start thinking of the total cost of ownership. That's a big deal for people for companies actually and a lot is wrapped up in that. At the end of the day though the biggest hurdle that most companies have when deciding whether or not to build or buy a system is that there needs to be a stakeholder who is going to really be responsible for it.
So generally in a buying scenario, and you see this a lot for big organizations and government entities the smart play is to understand all the features that your company needs. It could be a big system or a little system, you should come up with a matrix of sorts or a checklist. You want really good and clean user interfaces and user experience, an easy example is Google.
You also want to think of the ability to change the data model or AI model, and whether that's baked into the app and you know, rigid or whether you can actually affect those models that could and should be important to you.
Think that technology does 85% of what you want, and you're okay with that, but you can create a roadmap to get to closer to a hundred percent down the future, not costing anything. Finally, you want to think of ecosystems, like communities, the number of partners they have all the users, are people writing about that. Are there examples and forms and stack overflow in all these places do they have a large list of clients and do they provide training? Is there self-led training that they provide? Do they come in the house and do some training? So there's all of that for thinking of future proof in your purchase.
We've been talking a lot about our witness prep AI and what's gone into making all of that and today we thought we'd actually culminate into its intended uses and a bit about how we deliver it.
It's possible to acquire all the data that you need especially about cases and laws and decisions and whatnot you have to build a language model. There's been an academic language model called FILAC and we've extended that.
So given the ability to now understand legal texts and to model how to get to an outcome, basically to have all of them, I can render a decision, a law an actual case, whatever into a computer, a usable bit of data so that the computer can actually ingest it properly and not just have the unstructured text.
So now that you have this AI model, how do you deliver that to the best? You have to think about what the most natural thing for people is, when you're doing a user interface exercise you don't want to take people to a place where they have to think about the UI while they're also thinking about the context and problem that you're working on.
Because then your brain kind of fractures itself and you're lost of it. So a very common use case for most people just searches for a really good job of search and curation, which is something that we do.
It's all in the end powered by the AI model. But the way to deliver that is a very curated search especially if you're able to do it in a hierarchical view of what you're searching for so that you have, like a meta topic and you get yourself all the way down to specifics that get people the idea.
Most law offices & legal professionals go through common or similar cases, when they work on a case they look for prior arguments, prior cases, and precedents, allowing an automatization process to be implemented since it makes it searchable and usable information.
They use a sort of search tool and it's generally only a keyword search tool, since someone with a lot of experience, will know that in each case what to look at, the specific set of keywords that is pretty unique and doesn't usually happen in the rest of them. Of course, that other hand has some trouble since limiting to only those keywords you could lose relevant other information.
So when you need to know the exact keywords in order to search for precedent and prior cases. Imagine that there is the need to understand all of these details and information to really build a case up so that it's possible to understand what people have said in the past, what's worked and what hasn't, these being the material facts of the case.
The other thing that you really can't do without actual live debates and trials is to judge the strength of the argument, so this type of judging will obviously going to rely on the entirety of that package. In order to really understand and get to a point where there is confidence in an argument it really benefits everybody to, have sort of a score about it, so when there is an argument with a high degree of confidence, that there is the conviction that will or power up that case over the other side.
There's an opportunity here to make it more scientific and that's where we shine quite a lot in getting into the AI for legal decisions of courts and obviously also pushing the envelope and innovation in the world of witness readiness and witness prep.
In order to get to a world where you can build an AI program you're going to have to build up your internal capabilities and knowledge because you're dealing largely with unstructured texts, for example, images and video, those are definitely unstructured, but you can structure them by pre-processing in text. You can do the same, you can turn extreme, extremely large pieces of content, whether they are legal decisions.
We're still on the path of talking about legal, judicial, and general law with specific AI.
Today we're going to talk a little bit about the work of having a machine understand some of this law.
Clause extraction is often just looking for a way to segment certain documents, whatever clauses are you extracting from and it has many methods of extracting those. A clause is a self-contained paragraph or a section of text or content on a page, generally, it's not incredibly hard to pull them out. The difficulty starts happening a little bit more when you're trying to understand what those clauses are about, sort of classifying them.
There are a lot of possible classifications of these paragraphs or clauses which makes it like it's that next step beyond just separating them out. If you get into the legal domain or the legal world, basically the language inherently in those clauses is obviously going to be a little more technical.
So there are more caveats and specific understanding required of your machine learning or NLP programs to actually extract. Pulling the paragraphs out is not that hard, but once you have them out, it's again, understanding them from the legal perspective as to what they're for.
With this in mind, the FILAC Model was developed as a constraint to a specific domain of law to help structure the information.
There are other techniques especially in neural networks, in deep learning where without supervision so an unsupervised model, you can cluster information, helping group the different clauses and the intent of those in a way that suits your workload better.
So that sort of wraps up us talking about our language models a little bit, talking about clause identification, extraction and classification, and the whole building up of legal AI.
Did you know companies can manage their legal documents using NLP and AI?
So imagine that you have all kinds of PDF docs, some of which may have been scanned, so after that, you need to understand the scan itself, which takes time as well.
There is a technology called Optical Character Recognition, OCR and that's often used to get that data in and turn it into printed documents, basically understandable computer documents.
You now have text, which is a big important piece, but you need to understand the content. Your NLP and AI program are gonna have to do some kind of whizzbang magic to basically understand what's going on in all of that content.
Most good programs will generate internal metadata about the document or the content itself, it might even go down to the bolded words being there for emphasis and the italicized words being there for another reason.
That can also lead to an adjacent to summarizing, so you might wanna summarize the document so that you don't have this a hundred-page. The idea is to get the general concept of the whole document.
Today we're gonna look to talk a bit about the life cycle of AI projects and how they're undertaken, whether that's an internal or external project because they can both go well or poorly.
What we found was unless certain things are in place at your organization or with your people staff expertise, really experience, a lot of internal AI projects kind of go awry and don't exactly get to the outcome that's always desired or that was expected.
The quality of data relates to whether or not an internal AI project will succeed so we're kind of holding everything else aside and saying how do we isolate the data and ensure that the data is not the problem, cause that's really what we're trying to get to here.
So you wanna make sure that you have very complete data in that sense because having just a little bit of the signal and only understanding.
Missing data is going to lead to certain things weren't observed, and certain things weren't saved or cataloged, or curated, so you don't have a full picture.
You only have a signal or two, at this point is kind of like a needle in a haystack, if you don't have a complete data set, and of course, just generally bad data is not gonna help you.
So you wanna make sure that you have data in a way that's going to be usable down the line, clean data, so doesn't have all kinds of weird noise and bad signals in it, and you wanna make sure that labeled this data well.
Another problem with all the silo data is we often end up with duplicate data sets. So we may be describing the same thing in slightly different ways, but across the organization, you will duplicate the effort, wasting people's time, and having systems that we're probably paying for, for no reason, that is saving the same thing.
Today we're gonna focus on the interplay of ERP with AI and NLP, what 1000ML does and hopefully get to a place where we can understand how you can actually benefit from each other, and at the end how organizations can get a real return of investment.
ERP, Enterprise Resource Planning, and the main objective are cataloging the types of things that are in your organization.
The crucial things are the company resources such as people, places, or things and then also the material invoices, purchase orders, contracts, and all that stuff along with your inventory. The idea's sort of a centralized brain of all the things that make your business run.
ERPs capture and understand complex workflows and processes with all your resources, so, AI by itself is a largely bits and bites computational engine, or I guess inference and prediction engine doesn't do well with unstructured text. You have to find a way to make those documents into something that the AI can actually understand.
So that's where the entry point of NLP sort of exists here, and NLP gives AI the ability to source the information from those documents.
Today, we figured we would talk a little bit, but the considerations of when you are accountable or responsible for these projects themselves, and you want to ensure that you have the best return on investments.
We touched last week about all the things of an executive's train of thought that would go into such a project in an NLP and AI. The executive isn't necessarily going to dig into the implementation of NLP projects and they're not really going to source data for you, the executive wants to rationalize in their mind, do we have sufficient and the correct kind of data in this organization, or can we source it?
There are ways that you can actually look to get data and look to acquire data points that you may not have. Either you, create data and create metadata as needed, or you actually buy them from some catalog logs or databases.
If we're talking about NLP projects, we're often largely talking about unstructured data, largely it's written data, it's type data obviously, but it doesn't have the structure of, which makes it, that you have to create the structure from it. So that's often where the executive's mind needs to be, are they need to ensure that the enterprise or organization has enough NLP expertise in its staff, obviously that they can properly manipulate all this text.
It's like a mini-layer cake. It takes some work and it takes some foresight, but it's definitely doable. And it's possible, to think about this as an in-house project for sure.
Often, it also takes ingenuity and foresight, and thinking on the cultural level. So, you want to be an organization that is thinking about, innovative things and also considering new kinds of algorithms and new kinds of methods to do things and not kind of just stuck in what's working and what's paying the bills at all times.
We've talked a lot about AI, data science, NLP generally. And we've talked a little bit about some use cases, but today let's actually hit home with the payers of NLP. Today, we're very specifically and deliberately going to be talking about how NLP affects your business from an executive's point of view.
Regarding NLP, you get to a place where you understand that the content of language needs to be diced up and made digital in some way so that you can use computers, especially AI, to get some outcome, whatever that outcome is. So why don't we just hit home with the general idea of NLP because that probably informs what staff is doing?
So let's say you already have the content or you're putting in some documents of some sort that you've received, the NLP programs and programmers know the kind of details you need to extract, they'll do most of that work for you.
How would an organization go about implementing an NLP system? So if we take, you know, the decision out of the way that you're going to buy versus do it yourself, well, we may look at that later. But in, in general terms, your staff is going to look to ingest a lot of content. Now, you may already own and have that content.
Your time to deploy is usually a bit shorter and your cost may be about equal or not quite as high in the internal version. However, your support cost, your ongoing support costs. If this is not a key part of your business, can be too much to undertake as opposed to getting somebody else to support it.
So those are some of the considerations that you would have in implementing the NLP system at your organization. We talked a little bit about what NLP is to an executive, how it gets done at your organization, and the interplay between buying versus having your own staff do it.
Today we're going to talk to you guys about the myth that AI is here to kill you, and it's all going to result in this big singularity where AI gets so intelligent and takes over the world.
It's also not here to take over the world, it's taking over a few tasks. It's taken over a lot of computation actually sees it as mainstream.
Some of the earliest uses that companies and really companies and organizations started seeing with AI were really around, intrusion detection, bot detection, spam detection, so things that the company may not have been really aware of that they even had AI in their organization.
It's important to understand the general use patterns, of your websites, the applications of your mail, and everything, and you want to keep things around those metrics.
Another really common use has been chatbots, conversational AI, IVR so sort of like that first interaction with a company where you're kind of getting some information about them or talking a little bit about your problem with them. So it helps in sales and it helps in customer retention and in customer service.
When a company is looking to acquire a chatbot, they usually want the implementation to cost of minimum for a super bear bones chatbot, that's just capable of answering some questions, but that brings some problems such as the configuration, where you're going back and forth, trying to really discover what your intended audience wants to talk about and how to have all the question and answers.
So there's a lot of tinkering and refining that goes on and on and on for a lot of long time weeks since chatbots have an issue with multilingual support just simply because you're the one configuring the chatbot itself, in different languages for different purposes.
The world of chatbots is kind of pervasive if you have a cell phone or almost any television service. A version of chatbots or chat technology that existed has been evolving more and more over time.
So along with that came or after that came, natural language processing, which actually gives us the ability to turn language into more machine-readable and usable data, a bot can generate speech based on some ideas. So it understands the topics or the kind of answer that you would want because maybe this is strictly a customer service bot and it's transactional.
In today's episode, you can understand more about the Chatbot evolution and how they work.
Most companies, large or small, work according to written information from written documents that contain the details necessary for different operations. This wonderful amount of data has patterns and trends. It is possible to understand them in order to make an informed decision. The ability to not be tired and mess something up is huge, NLP can basically help take the guesswork out of your hands.
Talking about on employee productivity , if we replaced certain tasks which can be automated, a lot of what NLP really does for employees and just generally is find the needle in the haystack.
Most people prefer to do a bunch of research by themselves before they enter the sales cycle. If they ever enter the sales cycle at all, because they want to be informed and they don't wanna have buyers remorse though. This can sometimes also lead to buyers remorse if they don't go in there informed as well.
It has become universal for any kind of business to explore software development solutions that are effortlessly efficient yet unique, and gets the job done. Today, companies after companies are discovering no-code development and are embracing no-code platforms as a shortcut and expediting the process of their app development with little to no coding at all.
Let's find more about Low-Code and No-Code Development Environments.
Everyone has seen or heard of superintelligent AI, usually in the form of menacing robots. Luckily, those are mostly great fantasies by imaginative writers rather than a reflection of current reality. Today we venture into the things that AI can actually create. From the lowly meme to writing code and possibly creating other AI, let's find out how intelligent AI can actually (currently) be.
So much of law is opinion based that you may think that there is no way that AI could do an adequate job of automating some of the work. Fortunately, the legal and judiciary practices are very paper (document ) heavy, which lends itself superbly to NLP with AI. Today we'll explore exactly how the use of AI and NLP has changed the legal world and what may be to come.
Recently Data for Good and 1000ml decided to combine their expertise to respond to a sustainable housing challenge driven by the Canadian government. The innovative solution was put together with Data for Good volunteers with the help of AI experts at 1000ml and was delivered in 6 weeks.
Natural Language Processing with the help of Machine Learning is the current win-win combination used to detect fraud and misinterpreted information. One of the biggest challenges of the free and anonymous internet that we have constant access to and basically drives our life is “Fraud”.
How can you create efficient supply chain management? This is an open question for many suppliers, distributors, manufacturers, and retailers. Today, amid shifting supply chain market dynamics, changing ways of working, increasingly volatile demand, businesses are wondering how to make their supply chain less vulnerable to disruption. Machine learning holds the answer to many well-known as well as emerging supply chain challenges.
Contracting is a common activity, but it is one that few companies do efficiently or effectively. In fact, it has been estimated that inefficient contracting causes firms to lose between 5% to 40% of the value on a given deal. Recent technological developments like artificial intelligence (AI) are now helping companies overcome many of the challenges to contracting.
It seems as if the AI can be elusive when you don't have millions to spend on the type of talent that your small business would need to achieve big milestones. Fortunately, there are a myriad of ways in which your smaller organization can use AI; from customer service right through to sophisticated Document AI and Contract Management Systems, and much much more.
Organizations are constantly looking for an edge to win out against competition and the current state-of-the-art is to use AI. Becoming an AI organization is not for the faint of heart; let us guide you as to how you could quickly, safely and easily get your organization into the age of AI.
That's right, All Things Data is back!
Of all the possible sources of business value that exist in your organization, the documents themselves tend to be the forgotten clue.
With the amazing developments in machine learning, AI, and NLP, the time is now to gain a true understanding of your business through its documents, agreements, invoices, emails, intranets, etc... Really, any content whatsover.
How much does it cost your team in hours and dollars when one of your data jobs fails and you're running a report with inaccurate data? How do you REALLY know if all your daily data has been loaded properly? How long would it take to audit and fix?
Simply put, what happens when your data breaks?
This week on All Things Data we speak with Barr Moses the Founder and CEO of Monte Carlo. They've just closed a $16m round to solve data reliability through the concept of data down time.
Barr walks us through her time at the Israel Defense Force and her journey to the Bay area through Standford, Bain,working her way up through Gainsight and starting a company.
This week we'll cover:
What the inspiration was to start Monte Carlo
How the IDF help shaped her thinking
Why is data reliability important?
What is data downtime and why now?
Leveraging customer success as a super power
- Hubspot's data jingle!
Clearbanc is making waves in the VC world with their novel funding model and use of machine learning. We thought it was interesting to speak to Susan Shu Chang to see how data science is applied in the startup funding world.
We'll also be chatting about her journey from an Economics/Math major to RL (reinforcement learning) through her passion for gaming and how time management and prioritization has led her to get on the speaking circuit, writing a book, teaching, continue gaming and learning all while working full time!
If you're interested in how digitally native companies use marketing analytics at scale, this episode is for you.
Peter Cheung, the Director of Performance Marketing and Analytics at Mejuri stops by this week to talk about how Mejuri's marketing team uses data to continually grow their community of fiercely loyal customers during COVID-19 and beyond.
Peter gives us a peek under the hood on what he's looking for in a data marketer, why insourcing is right for them and how he thinks a data marketing team should be built.
Sr Data Scientist at Tucows, Kenny Kwan joins us this week to chat about why it's best to be a generalist when you're starting your data career.
This week we'll cover why it's important to be able to get data yourself, what does it mean to have "communication skills" (hint, it involves cannolis ) and how to advocate for yourself and your projects.
Kenny takes us through his story from leaving the Dominican Republic to pursue his PhD in the US and ending up as a data scientist in Toronto. Also, he's finally able become a Raptors fan this year because Kawhi isn't around.
Sr Data Scientist at Shopify, Fernando Nogueira joins us this week to chat about what it's like to work a large data driven organization and his path to getting there.
We go through what it's like to be the first at Freshbooks in their growth stage, ChatKit as the first data resource on a 5 person team and then transitioning to large data science group. Fernando gives us a peek under the hood at how Shopify uses data and is built to last. We also cover marketing analytics, why it's important to work on your development skills and how he almost becoming quant twice!
As an All Things Data Special Episode we combined forces with Dave Mathias from the Data Able Podcast to chat about Data Education.
We cover:
Where to start and where not to start your data journey
Where you should be focusing your learning
The spectrum of learning opportunities and how they fit into your data journey
How does your domain experience fit into the equation
Honing your craft with challenges like the 14 day Data Challenge from Tableau
Sr Machine Learning Engineering at Snap Travel, Joey Sham joins the All Things Data podcast to tell us what he's up to in his day to day as a machine learning engineer, how he made the transition from Physics to developer to data science, how Snap Travel uses NLP & BERT to help people find hotels and flights and how to hustle to get your first data science job.
Lead Data Scientist at Sunlife Financial, Jennifer Nguyen joins that all things data podcast to talk about how she uses data science in the insurance industry, how a web development project got her career started, what she's learning to become a better practitioner and advice for anyone looking to get in the field.
Apparently there are tens of thousands of Data/ML/AI jobs out there, but why isn't everyone you know getting hired and working in the field?
This week Jansen and Victor discuss:
How technology markets evolve
New talent vs mature talent (and the talent rush)
How exits and large companies drive innovation and talent pathways
What a transition looks like from school to the work force now and in the future
What are the things you need to consider from a data perspective when you build a product?
This week Victor and Jansen discuss some common pitfalls and risks they've seen over the years when building out products.
They'll cover:
Data first thinking
Data as your product
Building MVP's
Codeless development
When to bring data people to the table
We've been keeping all your DM's from LinkedIn, email and our meetup slack and complied it into our AMA.
This week we cover questions like:
What value does AI/ML bring to an organization?
I'm a CEO of a company, we need an AI strategy, our competitors are purporting to be AI driven. Where should I start?
What's the surrounding Hype of AI and what's the reality? - Why is there a race to AI now? Is it really something companies should be chasing
When do you think the data science hiring blitz will end?
I'm looking to break into the world of data science, should I be studying deep learning (sounds like a hot topic), or should I be studying the traditional ML techniques?
A data scientist is more than just scikit-learn and SQL.
This episode Jansen and Victor discuss the knowledge and tools that a data scientist should posses to be the real deal.
Why:
CI/CD
GIT
Containerization
Testing and more
Are important tools and concepts to have in order to be a successful practitioner.
How do you transform an organization that isn't digitally native?
This week Victor and Jansen discuss what it takes to enable legacy organizations to adopt AI.
They'll be unpacking:
Boardroom discussions
Executive sponsorship
Team structures
Project ideation
With COVID-19 still growing globally, physical distancing and contact tracing is our best bet to reducing the spread until a vaccine or treatment is available.
This week Victor and Jansen dive in to contact tracing tech, how it works and the adoption curve to make it impactful.
We'll be covering:
The Apple / Google tech
Bluetooth and wifi options
Other methods that may or may not be socially acceptable
Thoughts on how to maximize adoption
What does it take to become a global AI Super Power?
In this episode explore what countries are doing to win the AI race. Jansen and Victor unpack:
Education
Government policy
How the private sector plays in the equation
Which countries are investing and which aren't
What skills are needed and who needs them
Co-Operative education aka Co-Op.
What is it, why is it important and who are the players?
After speaking with leaders in education, industry and government for the last 3 months we thought we would share our experiences with helping people find their first job. Our journey led us to really deep into co-operative education.
What if you got measured on all the good and bad things you did in your day to day life?
How would we do it? Do we have the data?
This week Victor and Jansen delve into creating a social credit system. We answer:
What data would we need?
Is it a violation of privacy?
Who owns the data and who gets access to it?
How is China currently doing it?
We also reference The Information Trade: How Big Tech Conquers Countries, Challenges Our Rights, and Transforms Our World by Alexis Wichowski
Smart cities...
What makes them smart?
What's all the hype surrounding them?
Why are governments pushing hard for these initiatives.
This week Victor and Jansen unpack the city of the future. We'll cover:
Surveillance
Transportation
Smart Grids
Stories of smart city initiatives
Bonus points if you know what episode art is!
We've been keeping all your DM's from LinkedIn, email and our meetup slack and complied it into our first ever AMA.
This week we answers questions like:
What's next in technology?
How do you measure impact of a data science project?
What are our favourite thing about our jobs?
What should you do to make the next step in your career?
Our opinion AI making AI
Can data cure all?
Probably not.... but it definitely can boost outcomes. From cancer detection, chronic illness management and patient experience, data is the key to unlocking breakthroughs and service improvements across the sector.
This week Victor and Jansen discuss:
The social determinants of health
Wait times
Electronic health records
Connected hospitals
Drug discovery
Is retail as we know it dead?
COVID-19, Data, AI and Technology is accelerating change in the retail experience.
This week on the All Things Data Podcast Jansen and Victor discuss the future of retail:
Amazon Go
The rise of Phygital retail
Malls
Flu tracking?
New ways retailers are using your data for a better experience
When do I get my raise / promotion / secondment?
Find out data is changing the way companies evaluate, compensate and measure their employees.
(Hint: it's not the dreaded annual review)
This week Jansen and Victor discuss the shift in the way HR operates, the death of the annual review and how we measure workplace productivity.
This week we look at how data has changed the world of sports. From team building, player evaluation and gaining that competitive edge, data is helping professional sports team look at the game differently.
Victor and Jansen discuss the rise of the 3 pointer in the NBA, Moneyball in the MLB and F1 tuning to explore how a few numbers can change it all.
One of the big reason we're able to unlock AI is because of the amount of data we have in the world today. The Machine Learning models we build use this reference data to predict outcomes, this is why it's called training data.
This week Victor and Jansen unpack:
How we use data to predict outcomes
Caveats and watch outs about the data we choose
The cost of training and resources
Transfer learning and how we can leverage pretrained models
Critics worry that AI will take jobs away, but like all technology, it's about how we implement it. Will it be a world where it's humans vs machines or will it be humans + machines?
This week Victor and Jansen unpack the lanes, the tasks and uses cases humans and AI excel at. They'll also cover the implications on the job market and how AI will affect work as we know it (spoiler: we'll be OK).
As the saying goes, if it’s free, you are the product
This week Victor and Jansen discuss how your personal data is being used in the wild.
We'll be covering:
What data is being gathered
The tech behind data collection
What your data is being used for
Examples of the good, the bad and the downright scary uses of your personal data
The world is on lock down with the current COVID-19 crisis....
There is a lot of data floating around from different organization, governments and news which is hard to interpret, combine and answer questions.
This is your chance to make a difference as a data practitioner!
This week Victor and Jansen discuss:
The Data For Good not for profit organization
The inner workings of a data-thon and what you need to start your own
Types of data projects you could do to help your community
We've got available data teams at 1000ML to help, reach out and let us know how to.
In this week's episode we tackle the subject of What is required to land your first Data job.
Victor interviews Jansen on the intricacies of what it is like to start off in the world of data and try to get work in an environment that is radically changing as a result of remote work and now Covid-19
Part of our business is largely built on getting people employed in the world of Data. Most of our candidates we have come from a world of knowing how to study, take tests and write exams; in a very discrete fashion which does not lend itself well to the world of work. So we wanted to acknowledge and discuss what we are observing currently in the data hiring world and what employers like us are looking for in candidates.
Data Science is more than just Jupyter notebooks, it's:
Machine Learning... It's the ML in 1000ML.
Data Engeineering
Model Tuning
Infrastructure
And a whole lot more
This week Victor and Jansen discuss:
What is a model?
Data *'s aka Data People
Life beyond Notebooks
Data infrastructure
What it takes to productionize a model.
The business value of data science
This episode we'll be unpacking some the of the hype around AI.
We'll be covering:
Terms (and the marketing take over) like machine learning, deep learning and artificial intelligence.
Do you need a PhD to work in the data field
Where do these projects start at a company
The glue people and the foundation of a data team
Data unicorns, should you do it?
Our inaugural episode of the All Things Data Podcast.
Hiring and retaining top talent is hard! We'll be covering hiring and retaining data talent in a super competitive market and our experiences growing out some of the world's top data teams.