Written by text-to-speech experts; read by advanced AI voices. Explore the world of language technology and spoken-word audio publishing with news, insights, and thoughts from BeyondWords. Want to convert your written content into audio for your website, podcast platforms, and more? Get in touch.
Text-to-speech (TTS) is the artificial production of human speech based on input text. TTS can be viewed as a sequence-to-sequence mapping where text gets transformed into one of its many possible speech forms. However, this process is not as straightforward as it might seem. Any user of early sat nav systems should be able to recall numerous instances when the technology mishandled pronunciations, often resulting in an embarrassing slip. Nuances in pronunciation become particularly important when dealing with delicate news material or financial research. The main challenge of perfecting TTS lies in the nature of text itself, as it is inherently ambiguous and under-specified with regards to many aspects of speech. In this post, we outline the key challenges presented in processing input text and describe how a Text Preprocessor can help. Ambiguity in written language We can roughly categorise text into 1) natural language and 2) non-standard words (NSWs). Loosely speaking, natural language is any text that can be read out loud immediately, whereas NSWs have to be verbalised first. For example “April first” can be read out loud as is, whereas “Apr 1st” has to be converted into its natural language form first, before it can be read out loud, e.g. “April first” or “first of April”. There are many types of NSWs that need conversion, such as dates (“01/01/2001” into “January the first two thousand and one”), currencies (“$100” into “one hundred dollars”) or units (“10m” into “10 metres” or “10 million”), just to name a few. But even text without NSWs might require some form of disambiguation. For example, abbreviations need to be expanded into their natural form, like “St.” into “Street” or “Saint”. Another example are acronyms and letter sequences which can be read as words (such as “NASA”) or by pronouncing each letter separately (such as “FBI”). The process of disambiguating and expanding natural language and NSWs is commonly referred to as Text Normalisation. The output of which can be interpreted as a sequence of graphemes - letters that represent sounds in a written language. Ambiguity with respect to pronunciation The process of normalising text into a sequence of graphemes removes some of the ambiguity, albeit not fully. Take for example heteronyms, i.e. words which are spelled the same but pronounced differently depending on the context they appear in. For example, “lead” can refer to being in charge, in which case it's pronounced like ”leed”. However, “lead” can also refer to the metal, in which case it is pronounced like “led”. Pronunciation can also vary based on regional or personal preferences. For instance, the word “either” can be pronounced as “ee-thur” or “eye-thur”. Or consider the brand name “Nike”, which in the US is commonly pronounced “nai-kee”, whereas in other parts of the world people might say “nyk”. Furthermore, unknown words can appear over time, such as “Gif” a few decades ago. Some people use a hard “g” (like “gift”), while others use a soft “g” (like “giraffe"). Sometimes these new words deviate from existing orthographic rules, which makes inferring their pronunciation challenging. The examples above motivate the need to better specify the intended pronunciation of input text. This is commonly achieved with Phonemic Transcription, a process which transcribes the grapheme sequence into a sequence of phonemes. Phonemes are the smallest units of sound that can distinguish one word from another, and are represented by special alphabets such as the International Phonetic Alphabet (IPA). With hundreds of symbols it allows for more granular control than mere grapheme-based input. Ambiguity with respect to various aspects of speech Even after Text Normalisation and Phonemic Transcription, a reader may often have little information about many of the aspects of speech, such as timbre, which characterises who the speaker is (e.g. their age) or where they come from (e.g. their regional accent or dialect). It...
You package financial research into crisp and concise reports because your clients time is precious. For the same reason, you should make your insights available in audio. Listening enables engagement and multitasking where reading doesn't. Clients can consume audio briefings while they're driving to work or grabbing lunch, for example. They’ll have more opportunity to consume your valuable research and get the most from your products. It's also the practical option when they're away from their computer. Who wants to navigate a PDF on their phone, with awkward zooming and endless scrolling, when they can put in their earbuds instead? They'll no doubt appreciate the much-needed screen break. Convenience is particularly important when you consider the limited timespan of most financial research. If clients don't get round to your daily briefing today, they're unlikely to revisit it tomorrow. So, you really need to capitalize on every engagement opportunity. Audio can also improve accessibility, and many simply prefer listening or audio-assisted reading. Approximately 30% of Americans are auditory learners, and 60% of listeners say they process audio information more efficiently. So, as well as being more willing and able to engage in the first place, listeners may extract more value from your research. That will translate into better retention and conversion. The financial audio landscape Many financial services firms are already tapping into the demand for audio insights, using the format to demonstrate their digital-first approach and better engage target audiences. Since 2018, JPMorgan has delivered its annual letter to shareholders in audio. Not only is the additional format user-friendly, but its Chairman and CEO's voice lends the message personality and authenticity. The company also publishes audio commentaries from Global Liquidity Portfolio Managers on its website. UBS, Morgan Stanley, and Goldman Sachs are just some of the major players reaching listeners through podcast platforms like Spotify and Apple Podcasts. And their presence is attracting even more finance professionals to these apps, to the potential benefit of the entire finance industry. Meanwhile, publications like The Wall Street Journal and Bloomberg are using voice AI to deliver audio articles at scale on their websites and apps. Bain and Coronation are two of the financial services firms using BeyondWords text-to-speech to make their expert analysis and insights audible. The practical approach The case for audio hasn't only been strengthened by widespread adoption of smartphones and digital listening habits: advancements in voice AI have significantly improved the ease, cost, and speed of implementation. Using our API, WordPress, or Eidosmedia integration in combination with our player SDKs, financial services firms can make insights on their website or app listenable automatically. Audio is typically embedded within just a few minutes of publishing and dynamically updated. Audio reports and briefings can be shared via URL too, ideal for emails and PDFs. We also recommend that you auto-distribute via podcast feed to tap into existing audio habits. Subscribers on platforms like Spotify and Apple Podcasts will receive the latest episodes direct to their app, with notifications prompting engagement. You can even build custom feeds or playlists to group audio by category. Research teams also have the option to create or edit audio in our Text-to-Speech Editor. Intros, outros, and sound effects are easily inserted to enhance the listening experience and strengthen sonic branding. Firms get access to our AI voice library, but may prefer to create and use voice clones of their financial analysts or other stakeholders. This can lend the briefings more credibility and assist with relationship-building, without additional demand on already-busy staff. All voices are supported by our customizable natural language processing algorithms, which ensure accurate and ...
The Irish Times is the most popular digital news subscription in Ireland, with 39% of all subscribers saying they're signed up. It's also one of the most valued and trusted news brands in the country. One of the many features contributing to its success? Audio. We spoke to Paddy Logue, Digital Editor at The Irish Times, about the publication's adoption of voice AI and its role in a wider news audio strategy. The move into audio The Irish Times began offering audio articles with BeyondWords in 2019, when listening habits were really picking up pace in the country. Earlier that year, Reuters said that podcasts appear to be reaching critical mass and found that 37% of Irish people listened to podcasts monthly. Podcasts remain extremely popular in Ireland. Last year, the country led the pack among 20 international markets, with 46% of the online population listening monthly. Logue said, quote. We can see, through the growth in podcasts, audio becomes more of a typical activity for news audiences who don’t necessarily have the time or inclination to sit down and reed. Recognizing that our audience was often on the move or busy with other things in life, while at home, we wanted to provide an alternative or complementary way for them to consume our journalism. We very much believe that what is most important is the journalism itself and the means of delivering it, the platform, should be influenced heavily by the audience preference. End quote. At the same time, voice AI had reached a tipping point. With the development of lifelike AI voices and automated audio publishing tools, there was finally a viable alternative to human voice over. The Irish Times knew they needed to take control of their audio UX and sonic branding. Logue said, quote. We were very impressed with the quality of the AI voices, and the customer service and support during implementation was seamless from the BeyondWords team. End quote. Using audio to drive subscriptions In Ireland, 16% of the population pays for online news. Subscriber only features, like the Listen audio service, are helping The Irish Times to capture the largest share of this market. The Listen brand brings together various podcasts and audio versions of articles. And while a selection of opinion columns and feature articles are author red, with the audio hosted on BeyondWords, voice AI is crucial to delivering the bulk of news audio on time and at scale. Integration with their Arc XP content management system, alongside our player SDK, means audio versions are created and published automatically. Logue said, quote. We now have AI audio on almost all of our articles and we have decided to make this a subscriber benefit. When a user clicks the icon to listen they are prompted to log in or subscribe. We see it as a major part of a more holistic audio offering. End quote. And it's not just about enticing people with listening needs and preferences to sign up. Another key benefit of audio articles is reduced churn. When subscribers have the option to listen — not just reed — they're better able to fit news into their day. That means they're less likely to experience unread guilt factor, the reason behind 13% of digital news subscription cancellations. Plus, Northwestern University research shows that frequency of consumption is the biggest predictor of retention in digital news. A localized voice for a localized audience The Irish Times saw positive results with other voices, but it was the adoption of Irish accented voices that really cemented their commitment to automated audio. The availability of an Irish accent was a big determining factor in our expansion in the use of AI on our newly designed site, says Logue. Not only does the localized voice allow them to better connect with their localized listeners, but it delivers more accurate pronunciations of regional words. This works in conjunction with voice support from BeyondWords. For example, thanks to our customizable NLP algorithms, we'r...
Today, we've added four premium voices to our AI voice library — Jodi, Adam, Kate, and Maria — all created in collaboration with professional voice actors. This means it's easier than ever for American English publishers to find an AI voice that resonates with their listeners. The voice you're listening too now is Jodi. Hello, this is the voice named Adam Hello, this is the voice named Kate Hi there, this is the voice named Maria These lifelike AI voices are ideal for a variety of audio projects, from audio articles and newsletters to listenable landing pages. They're available exclusively to paying users under terms set by the voice actors themselves — just drop us an email on hello@beyondwords.io if you'd like to request access through our text-to-speech platform. How were these voices created? We believe that AI voice actors should be fairly compensated and retain control over their voice IP. That's why every partnership starts with our Voice Cloning Contract. This open legal template allows voice actors to license their voice in a way that's fair, safe, and rewarding. Next, the voice actors recorded scripts specially designed to get the best speech data. Our text-to-speech engineers then cleaned up the recordings and trained the voice models, putting them through various rounds of testing and improvement until they met our and the voice actors' exacting standards. When we asked Kate Marcin, the voice actor behind Kate, about her experience creating the voice clone, she said I thought it was great, very easy. I narrate a lot of audiobooks and do a lot of other long-form narration so it was the same type of work for me. Maria Pendolino, the voice actor behind Maria, said I love knowing that my voice work can help others experience content in a way that makes it more accessible and available to them. I'm excited to offer this as another part of my voice over business and performance. To request access to these premium voices, email hello@beyondwords.io
Today, we're launching a US version of our Voice Cloning Contract or VCC, formerly the Voice Services Agreement Template or VSAT, in partnership with the Open Voice Network. This open legal template facilitates safe contracting between voice actors and companies creating AI voice clones. The VCC empowers voice actors to license their voice IP under a set of transparent terms. The contract aims to give voice talent confidence while working in AI and lays out how royalties should be returned on a usage basis. Voice actors can also specify in which sectors or use cases they would have their clone used. We have made the US and UK versions of the VCC freely available for anyone to download and use. It's part of our commitment to making the AI voice industry safer, because we believe that a fair approach benefits everyone. BeyondWords has already used the VCC to create voice clones of UK and US voice actors. These premium AI voices are available exclusively to BeyondWords subscribers that meet the actors' terms. As well as earning an upfront recording fee, these performers receive royalties based on the amount their voice clone is used throughout the term of the contract. This means that they earn passive income while still working on traditional voice over projects. Voice actor Kate Marcin said. I decided to work with BeyondWords because I realized that it was inevitable that this technology would be developed and continue getting better. Rather than trying to fight the inevitable and be worried about my career as a voice actor, I wanted to work with a company that actually has voice actors' interests in mind and wants to work directly with us rather than stealing voices as I've heard some other companies have done. Voice actor Maria Pendolino said. I love knowing that my voice work can help others experience content in a way that makes it more accessible and available to them. I'm excited to offer this as another part of my voice over business and performance. Interested in licensing your voice to BeyondWords? Register your interest using the form below.
Giving newsletter subscribers the option to listen — not just reed — is an effective way to improve engagement and retention. And providing audio versions of your emails is easier than you think. Keep reading or listening to learn the benefits of making newsletters listenable, how to create audio versions of your newsletters, and how to add audio to emails. If you're a Substack writer, check out our article'Why and how to add voiceover to Substack posts' instead. The benefits of making newsletters listenable When subscribers are too busy to read or just not in the mood too, they'll probably delete your newsletter or just leave it unread. If that becomes a habit, they're likely to unsubscribe. Providing an audio version gives them an alternative and allows them to consume your content while they're on the go. Tapping into these listening needs and preferences can therefore improve newsletter engagement and reduce subscription churn. Audio functionality can also help you attract more subscribers in the first place, because spoken-word audio content is more popular than ever. With 71% of listeners saying that they multitask, 60% saying that they process audio information more efficiently, and 56% saying they prefer listening to reading, it's clear that listeners will consume content they'd never even consider in a written format. Another benefit of making your emails listenable is that you can distribute the audio newsletter independently. This allows you to expand your reach through platforms like Apple Podcasts and Spotify, or draw attention to your newsletter offering through on-site playlists. You can even monetize your audio newsletters with audio ads. How to create audio versions of your newsletters You can create audio versions of your email newsletters using human voiceover or voice AI. If you have the right tools, equipment, and skills at your disposal, recording the narration yourself or hiring a voice actor can give you an edge on quality. However, this can be time-consuming and expensive, making it difficult to meet your publishing deadlines and achieve a positive return on investment. Voice AI is the more practical option, because you can produce engaging audio versions affordably and in a matter of minutes. With our Text-to-Speech Editor, it's just a case of pasting your text and clicking 'Process audio'. You can even automate audio production if you're using a CMS like Ghost. We've got a wide range of advanced AI voices to choose from, and it's possible to create a custom voice — ideal for publishers who want to strengthen their sonic branding, as well as writers who want to add a personal touch. Plus, you can elevate your audio versions with sound effects and other features. How to add audio to emails The way that you add audio to your newsletter will have a significant impact on engagement. You need it to draw the attention of would-be listeners without imposing on your readers' experience. Unfortunately, the majority of email clients don't support audio, so it's not possible to embed an interactive audio player like you can on a web page. You might be able to attach audio files, but this can result in your newsletters going to spam — and won't provide a good user experience anyway. In most cases, the best option is to link too an online audio player. With BeyondWords, you can just copy the audio's shareable URL and link to that. The web page features the customizable Large Player, which provides a great listening experience on desktop, tablet, and mobile. It also has analytics. You can attach your link to a button, image, or text. The image could be anything from a player mockup to a headphones icon. It's a good idea to A B test different designs to see what works best for your audience. If you're using an editorial newsletter platform in which your emails are also published as posts, it may be possible to embed an audio player into the web page. This will provide a good listening experience online, but bear ...
Audio articles are on the brink of ubiquity. There was rapid development in the format last year, according to Reuters Institute, and 80% of media leaders are planning to invest more in digital audio throughout 2022. Nic Newman wrote In our conversations around trends and predictions, it is clear that many publishers believe that audio offers better opportunities for both engagement and monetisation than they can get through similar investments in text or video. Let's take a look at some of the leading news publications making their stories listenable. The New York Times With 7.6 million subscribers as of September 2021, The New York Times is the world leader in digital news subscriptions. It's also a major force in audio journalism, reaching around 20 million listeners a month. The publication is currently testing New York Times Audio, an app that brings together audio journalism and storytelling, as well as narrated articles, podcasts and audio content from a slate of premier publishers. The Daily, its flagship podcast and audio newsletter, reaches around 4 million listeners every day — that's almost twice the print edition's peak circulation. It’s helpful in driving affinity to the brand. It’s harder to track directly how it drops people into the core news subscription funnel, but we have every reason to believe it does, given how our results have improved. What I would say on audio is that’s a place where we think there’s going to be real demand at high CPM for some time to come, and our product set is expanding. I think audio will be something to watch in our ad business for some time to come. Said Chief executive Meredith Kopit Levien. The Washington Post With 3 million subscribers as of November 2020, The Washington Post is the second biggest news subscription service worldwide. In 2020, readers who listened to audio articles on The Washington Post's apps engaged more than three times longer with their content. This persuaded the publication to adopt voice AI, so that they could deliver audio articles at a larger scale. We conducted user research and learned that users want to stay informed but are busy, so they appreciate an option to get up to speed on the latest news developments while cooking dinner, running errands or exercising Said Emily Chow, director of site product. What we’ve learned from users is that they listen to the news while doing other things, and are consuming far more content than they would normally. Said product manager Leila Siddique. The Irish Times The Irish Times is the most popular digital news subscription in Ireland, with 39% of all subscribers saying they're signed up. The publication's Listen brand plays an important role in the publication's revenue strategy, helping them to increase conversions and reduce churn. Using a BeyondWords integration with Arc XP, The Irish Times began offering automated audio articles in 2019. They now use an Irish accented voice from our library to deliver engaging narration on their website and app, as well as distributing their human-read audio through our platform. We now have AI audio on almost all of our articles and we have decided to make this a subscriber benefit. When a user clicks the icon to listen they are prompted to log in or subscribe. We see it as a major part of a more holistic audio offering. Said digital editor Paddy Logue. The Wall Street Journal With 2.8 million subscribers as of September 2021, The Wall Street Journal is the third biggest news subscription service globally. The publication has been pretty quiet regarding its audio strategy and results, but it delivers audio versions of most articles using AI voices. The Economist The first major publication to invest in an audio edition, The Economist has long flown the flag for listenable journalism. It all began when the publication identified a problem it called unread guilt factor. Subscribers were canceling because they didn't have enough time to read. The idea is just to leave i...
In 2011, The Onion published an article entitled: The Economist To Halt Production For Month To Let Readers Catch Up. It was satirical, but it probably touched a nerve at The Economist HQ. After all, a lack of reading time was the number-one reason for subscription churn. Five years later, the publication publicly acknowledged the article — and the problem. "The Economist can be overwhelming," wrote Product Manager Richard Holden. "It can be difficult for even a motivated reader to keep up." When subscribers watch print editions pile up, see headlines stay unclicked, they feel guilty. How can they justify paying a monthly fee for something they're barely using? Who needs a constant reminder that they're falling behind on their reading goals? And so, they cancel. It's a problem that The Economist's then–Head of Strategic Development, Denise Law, referred too as unread guilt factor. She wrote, quote. "Many of our former subscribers found it difficult to stay on top of The Economist edition each week and simply gave up. We spoke to some of them to validate that the problem actually existed. It did." End quote. Subscription guilt isn't unique to the news industry, but it is particularly prone too it. With such a huge volume of content published every day, it's easy for readers to get overwhelmed. Nieman Lab found that 13% of news subscription churn results primarily from people having too much to read in too little time. So, how are publications tackling subscription guilt? With audio. Tackling unread guilt factor with audio Giving subscribers the option to listen — not just reed — means there are more opportunities to stay updated. They can consume your content while they're driving, exercising, or cooking, for example. The Economist's Tom Standage said, quote. The idea is just to leave it up to the reader to decide what the most convenient form of consuming the content is, and in many situations that will be audio. We get a lot of feedback from people saying that it is how they stay on top of the information coming their way. End quote. In fact, many people go out of their way to multitask with audio: 25% of podcast listeners listen to fill empty time. While readable content is somewhat of a burden, listenable content is actively sought-after — something to keep minds occupied and entertained during mundane or physical tasks. Leila Siddique, the Washington Post’s Senior Product Manager, said, quote: "What we’ve learned from users is that they listen to the news while doing other things, and are consuming far more content than they would normally." End quote. The New York Times identifies these extra opportunities for engagement in what it calles the 'Audio Day'. Engage subscribers at these times, and they're far less likely to feel guilty and cancel. Plus, Northwestern University research shows frequency of consumption is the biggest predictor of retention in digital news. Standage said, quote. "Our evidence suggests that the audio edition is a very effective retention tool. Once you come to rely on it, you won’t unsubscribe." End quote. And let's not ignore the many other benefits of offering your articles in audio. It's not just time-poor readers who are turning to audio articles, audio newsletters, and podcasts. There are all kinds of reasons that people listen, and many ways to capitalize on audio habits. AI brings audio to the masses At first, The Economist's success was hard to replicate. The publication employed a large team to produce its audio editions, and they had an afternoon a week in which to do so. Those with fewer resources at their disposal, and shorter deadlines, found audio publishing infeasible. Our text-to-speech platform changes the game. We give publishers access to lifelike AI voices and a voice cloning service, and offer tools that allow for seamless integration into the publishing workflow. Engaging audio versions can be distributed effortlessly, affordably, and at scale. There is finally a viable al...
In the BeyondWords changelog for July 2022, you'll learn about creating custom text to speech rules with ayliasiz, reviewing audio before distributing with pending review, changing the position of audio ads, adding categories for google podcasts, new languages and voices, and fixes and improvements. Keep listening to learn more about what's new. Create custom text-to-speech rules with ayliuhsiz You can now add custom text to speech rules using ayliasiz. This is helpful for expanding abbreviations. For example, you can decide that M S F T is always pronounced as Microsoft or sen is always pronounced as senator. You can also provide pronunciations for sensational spellings. For example, music publications can ensure the band name spelled C H V R C H E S is correctly read aloud as Churches. Don't want the alias to apply every time? Open the Text-to-Speech Editor to add, edit, or delete ayliasiz in a particular audio. Let's say you're usually happy with it being pronounced NBA but on this occasion you're writing about the National Boxing Association rather than basketball. You can set an audio alias to prevent any ambiguity — and without having to make extensive edits to your text. Want more complex customizations? Submit a request to our team. Review audio before distributing with Pending review Want to review your audio before it's distributed? Our 'Pending review' feature makes it easy. When 'Pending review' is enabled, processed audio goes into an unpublished state by default. This means that you can proof-listen and make any edits before publishing your audio. For more advice, check out the audio publishing workflow ideas on our knowledge base. Change the position of audio ads You can now decide whether your audio ads are played pree roll, mid roll, or post roll. This can help you to boost engagement with your sponsored messages and audio. Just head to the 'Ad positioning' section in project settings to get started. Add categories for google podcasts If you've red or listened too our guide to optimizing your podcast feed, you'll know that choosing the right podcast categories can help your show to appear in relevant categories, charts, personalized recommendations, editorially curated collections, and search results. And we've now made it easy to optimize for Apple Podcasts, Spotify, and Google Podcasts. Just choose the relevant categories when setting up your feed. Remember, you can set up a podcast feed for an entire project or any custom playlist. Try new languages and voices Over the past three months, we've added a variety of new voices and languages too our AI voice library, so you might just find a voice that better resonates with your audience. Head to the settings to see what's available in your language locale. Fixes and improvements Support for Ghost 5.0 Faster reprocessing speeds Improved playlist S D K Improved docs and guides Eager to try out these new features? Sign in or sign up now at beyondwords.io to get started. You can upgrade your plan anytime from your BeyondWords dashboard.
Modern audiences are engaging with audio content like never before. Spoken-word audio's share of digital listening increased by 40% from 2014 to 2021. Almost half of over-13s in the US are listening for 2 hours and 6 minutes every day. Podcasts, audiobooks, and narrated content are more prevalent and popular than ever. Listening needs and preferences are nothing new, so what's behind this audio boom? It's something we call the audio flywheel. What is the audio flywheel? The flywheel effect is a business metaphor developed by Jim Collins. It suggests that multiple small efforts are required to build momentum, and that momentum will then perpetuate itself. We can apply this concept to audio content. Various factors have contributed to the rise of audio, and now that it has built momentum, its growth is showing no signs of slowing. Consumers, creators, advertisers, and technology companies all form part of the audio flywheel. Developments in one area have a knock-on effect in others, forming a positive feedback loop. Getting the wheel turning One of the main contributing factors to the rise of audio was increased accessibility of audio publishing and broadcasting. While the internet made it possible for anyone to share digital audio, hardware and software developments made production more affordable. This empowered more creators to explore the format's potential and attract listeners with new types of content, such as podcasts. Recently, advancements in text-to-speech have reduced the barriers even further. Our lifelike AI voices and sophisticated tools provide a practical alternative to human voiceover, allowing creators to produce engaging spoken-word audio quickly, easily, and affordably. This has opened the doors to previously unviable formats, such as audio articles. Meanwhile, ownership of audio-enabled hardware has rocketed. Smartphones and AirPods make it easy to listen anywhere, while bluetooth speakers and smart speakers promote listening in the home. Improved in-car connectivity means listening to your own audio while driving is just as convenient as switching on the radio. There has also been widespread adoption of audio apps like Apple Podcasts and Spotify, which are constantly developing new ways to keep listeners engaged. Changing media habits Consumers began adopting new media habits too. Most importantly, a proclivity for multitasking. Technology gives us 24/7 access to information, and our brains deliver dopamine when we learn something new or feel productive. So, we're prone to filling every minute of spare time with content consumption. Audio is the most multitasker-friendly form of media: we can listen while we're driving, exercising, cooking, cleaning, getting ready...the list is endless. Combine this with the increased availability and accessibility of audio content, and it's little wonder it's become the format of choice for busy people. 71% of monthly spoken-word audio listeners say the ability to multitask plays a role in their habit. Concerns around screen time are also starting to play more of a role in audio consumption. More than half of Americans are on their phone more than they'd like, and around a third are taking steps to curb their habit. As a result, more are turning to audio formats, which fulfill their desire for information and entertainment without gluing their eyes to a screen. Audio goes commercial As audiences spent more time with digital audio, advertisers came along with them. And they saw impressive results — one Nielsen study found audio ads delivered a 24% higher recall than display ads and were twice as likely to drive purchase intent. US podcast advertising revenues exceeded $1.4 billion in 2021, equivalent to 72% year-on-year growth — twice that of the total internet advertising market. The IAB said this was fueled by "a continually expanding user base consuming a growing library of engaging and diverse content". The subscription economy has also been picking up pace. The pro...
Modern audiences want content narration. And if you don't provide it, third-party tools will. It's already happening. The Read Aloud Chrome extension has over 4 million users alone. And as listening habits become more entrenched, we're likely to see more sophisticated screen-reading solutions come along — and higher uptake of the technology. This might sound like a good thing for publishers, like third-party tools handling all the hard work for you. But there's a big problem: losing control of your audio user experience, or audio UX. You don't get to decide what your audio sounds like. You can't optimize the listening experience. You don't get ownership over the audio content. This means missing out on valuable sonic branding, engagement, and monetization opportunities. Keeping creative control Whether they use text-to-speech or human narration, publishers who handle audio production themselves have full creative control over their audio UX. Unlike publishers who rely on user-side solutions, they can increase audio engagement by building custom players for their website or app, distributing their audio through other channels, and creating different types of audio content. They can also boost listener retention and strengthen their sonic branding by using the best voices for their content and target audience, ensuring their audio content is narrated accurately, and elevating audio versions with sound effects and other features. They can even measure their success through audio analytics, or directly monetize their listenership through audio ads and subscriptions. All of these efforts are about maximizing the benefits of audio publishing. But they also play a part in a wider branding and marketing strategy, helping publishers to foster a positive and pervasive image — and stand apart from their competitors. Berlingske, for example, has built a custom audio article player using our JavaScript SDK. This prompts users to sign up in order to listen, helping the publisher to drive new subscriptions and reduce churn. The team at Aftenposten have worked with us to create a voice clone of their most popular podcast host, which means their audio articles are read by a tried-and-trusted voice. They have also collaborated with our engineers to improve how the AI voice pronounces local and out-of-vocabulary words, ensuring the intricacies of the Norwegian language are handled accurately. This has led to a 75% increase in listening duration. Protecting your sonic brand In 2009, the Author's Guild successfully fought against default text-to-speech on the Kindle, to protect authors' and publishers' rights to control and profit from the audio versions of their books. Digital publishers need to follow in their footsteps. To seize ownership of their audio publishing rights. It's only a matter of time until audio narration becomes ubiquitous online. Publishers who act now can build their audio UX on their terms — start growing and capitalizing on their listenership. This puts them in a stronger position in today's media landscape and helps them futureproof for tomorrow's. BeyondWords is the practical solution. With our AI voices and CMS, you can publish branded audio content without undue effort or expense. Create your free account at beyondwords.io to explore the platform for yourself or email hello@beyondwords.io to request a demo.
An efficient workflow is essential in an increasingly competitive media industry. 67% of publishers say automation is of increased importance this year, and 81% of media leaders say that AI will be crucial for newsroom automation. It's little wonder then that those looking to capitalize on the benefits of audio publishing are turning to voice AI over human voiceover. There's no need for costly recording sessions or complex logistics — the technology can handle all the hard work for you. But using voice AI doesn't have to mean handing over the reins completely. A little human input can elevate text-to-speech audio to the next level, allowing you to strengthen your sonic branding and boost listener engagement. In this article, we'll help you find the audio publishing workflow that's right for you or your team. The stages of an audio workflow 1. Creating audio You can create audio automatically by connecting your CMS to BeyondWords. If you use WordPress or Ghost, our plugins make this very simple. Otherwise, you will need to use our RSS Feed Importer or API — this will require some development knowhow. Once setup is complete, newly published content will be auto-imported to BeyondWords, where it will be converted into speech using your default voices. You can control which elements are converted into audio, and how they are pronounced, with custom text-to-speech rules. By default, your audio will be automatically updated in line with changes to the source text. This means that the narration will always remain accurate. However, if you wish to differentiate the audio version from the text version, you can manually edit it any time. You can create audio manually using our Text-to-Speech Editor. This allows you to turn any text into speech using any combination of voices. You can also insert your own audio files. Whether you're repurposing written content or writing audio from scratch, you have the opportunity to optimize your text for an audio format. You can also use different voices and sounds to enhance the listening experience and strengthen your sonic branding. 2. Reviewing and editing audio If you would like to review your audio before it is distributed, you can enable the 'Pending review' feature. This means that audio goes into an 'Unpublished' state by default, and must be manually "published" before it can be shared with your audience via an audio player or podcast feed. While this adds an extra step to your publishing workflow, it does mean that you can perfect your audio before anyone else hears it. This keeps listeners listening for longer — and coming back for more. After downloading and proof listening to your audio, you can fix any problems by reporting a voice issue or adding a custom text-to-speech rule. You can also make edits in the Text-to-Speech Editor — edit your wording, change the voice associated with any paragraph, or insert your own audio files. If your audio was created automatically, you may wish to disable automatic updates, to prevent any changes from being overwritten via your CMS. Once you're happy with the audio, you can change its state to 'Published' via the dashboard or API. It will then be ready for automatic or manual distribution. 3. Distributing audio Automatic distributions help you to streamline your audio publishing workflow. They mean you can get your audio in front of your audience with no extra effort. If you are adding narration to web or app content — especially at scale — we recommend automatic embedding. This means that the audio player is automatically embedded into pages with audio versions, so that the page is made listenable. The customizable Small Player UI, specially designed for on-page narration, is used by default, but you can build your own player UI using our JavaScript, iOS, or Android SDKs. You can also auto-distribute audio via playlists. Just create a playlist with custom rules, and qualifying audios will be added automatically. You can distribute these playlists...
Converting articles into audio is one of the most rewarding audience engagement strategies. Not only can adding narration be quick, easy, and affordable, but people are listening more than ever. Keep reading or listening to learn about the 7 key benefits of converting articles into audio. 1. Improve engagement Making your articles available in audio allows you to better engage audiences with listening needs and preferences. For example, people who are multitasking, prefer listening to reading or reading and listening simultaneously, have certain vision or reading difficulties, or want to take a break from their screen are more likely to stick around when audio narration is available. Plus, studies suggest that audio can resonate more deeply than other content formats. U C L found that listening to audiobooks elicits a more intense physiological and emotional reaction than watching films or television, while the National Literacy Trust says that listening can aid comprehension. So, when people do opt to listen, they may have a more positive experience with your content. This can have a number of benefits, such as making them more likely to return and recommend your articles to others. 2. Increase traffic There are many cases where someone will want to listen to an article but not read it. If you cater to this audience and target them through your marketing efforts, you can drive more traffic to your site. You can promote audio functionality in social media posts, meta descriptions, and many other areas to secure click-through from listeners — not just readers. Plus, people may be more likely to share and link to your content if it has an audio version, because they understand this makes it accessible and engaging to a wider audience. This can boost your traffic even further. 3. Improve accessibility Adding audio narration to your articles improves accessibility and inclusivity for people with vision impairments and reading limitations. This includes people with dyslexia, blindness, and low literacy. Another benefit is that you have control over how your content sounds, because users don't need to resort to screen reading software, which can provide a poor listening experience. One of our publishers' readers, Troy Phillips, said: "Audio articles help me get the news easier and faster. I have dyslexia, which makes it hard to read the news. Screen-reading software is unreliable, and I don't always have access to it. So I really appreciated it when my chosen news source started using BeyondWords." 4. Increase subscriptions and reduce churn In a subscription-based revenue model, offering audio articles as a subscriber perk can improve your conversion rate. In the US, 36% of digital news subscribers say one of the main benefits is getting content that is only available to paying customers. In Mexico, access to more audiovisual content is one of the most important factors in digital media subscription decisions. And in the UK, 48% of digital news subscribers say they pay for online access because they get a convenient package of news and information. However, a bigger benefit may be reduced subscription churn. 13% of news subscription cancellations result from subscribers having too much to read in too little time. The Economist's then Head of Strategic Product Development, Denise Law, referred to this as "unread guilt factor" when talking about the publication's move into audio. When you offer audio, subscribers can consume articles while they're on the go. And because they find it easier to make time for your content, they're more likely to feel that their subscription is justified. Tom Standage, Deputy Editor at The Economist, said: Our evidence suggests that the audio edition is a very effective retention tool; once you come to rely on it, you won’t unsubscribe. 5. Increase ad revenue Audio articles can also increase revenue in advertising-based models. This is a knock-on effect of stronger audience and engagement metrics, whic...
With our new custom playlists feature, you can create and deliver multi-audio experiences. This opens up new and exciting ways to make your words heard. Using text-to-speech audio or uploaded audio, you can now: Create playlists manually Create a fixed list of audios that's editable anytime — perfect for splitting longer content into sections. A Creator, Pro, or Enterprise plan is required. Create playlists automatically Add custom rules to build playlists that are auto-updated with qualifying audios — perfect for grouping audio by theme. A Pro or Enterprise plan is required. Each custom playlist has its own Playlist Player, which can be embedded into your website or shared via URL. This provides an engaging listening experience, allowing listeners to easily switch between audios, change the playback speed, and more. It can also be customized with your chosen colors and image. You can even distribute your playlist as a podcast — just customize its podcast feed and submit the URL to platforms like Spotify and Apple Podcasts. This is a great option if you're publishing audio newsletters. Not sure what type of playlists to create? On our knowledge base, we've got 3 audio content playlist ideas to inspire you. Ready to get started? Already signed up? Head to the 'Content: Playlists)' section of your project dashboard to get started. You may need to upgrade your account first. New to BeyondWords? Create your free account at beyondwords.io and choose the pricing plan that works for you. Check out our docs and guides or get in touch if you have any questions.
Creating and sharing playlists is an effective way to improve engagement with audio content and narrated text. It aids with content discovery, helping you connect audiences with audio they're interested in — whether it's on your website, podcast platforms, or elsewhere. Here are 3 audio content playlist ideas to inspire you. 1. Split longer audios into sections Best for white papers, reports, and e-books. Users don't always want or need to listen to an audio in full. Perhaps they'd like to hear a specific section of your report, skip to a later step in your how too guide, or pick up from where they left off earlier. Splitting your audio into sections helps listeners get straight to the parts they're interested in. With our Playlist Player, users can listen right through everything at the touch of a button or jump straight to a specific chapter. That means a better user experience, which should translate into better engagement. Step 1. Create audio for each section in the Text-to-Speech Editor Step 2. Manually create a playlist featuring all your sections, making sure they're in the right order Step 3. Customize the Playlist Player, then embed or share via URL As well as embedding the playlist at the beginning of your content, I recommend embedding the individual audios into each section, using our customizable Small Player for a streamlined look. This enables listening from any point and makes it easier to listen and read simultaneously — especially if your content is spread across multiple pages. 2. Make targeted recommendations Best for news articles, blog posts, content marketing, and creative writing. Custom playlists make it easy to present related audios together. This means that you can curate audio recommendations for various purposes and audiences. Let's say you've published a story in three parts. As well as embedding the corresponding audio version into each page, you could embed a playlist featuring the entire series. That way, listeners won't have to navigate elsewhere to catch up on what they've missed — or hear what's next. When you create playlists automatically, you can deliver themed and ever-evolving recommendations. Just set up the rules, and the playlist will be auto updated with relevant audios. For example, news publishers can use meta data to group audio by category or author. These targeted playlists can then be embedded into associated landing pages and articles, or linked to in personalized newsletters. It's a near effortless way to help audiences find audio they'll like. 3. Manage audio newsletters with ease Best for news bulletins, daily briefs, and market insights. When busy audiences don't have time to read or listen to every piece of content, an audio newsletter is a great way to keep them informed and engaged. You can convert existing roundups and newsletters into audio or create your audio briefings from scratch. I recommend distributing your audio newsletter via podcast platforms like Apple Podcasts and Spotify. This makes it easier for podcast users to find and listen to your updates. They can even subscribe to your show to receive the latest episodes direct to their app. There are two ways to achieve this in BeyondWords: Option 1. Create newsletters in a separate project — recommended if you are creating the audio manually Option 2. Set up an automatic playlist for newsletters — recommended if you are creating the audio automatically You can then create a podcast feed for your project or playlist and submit it to the directories of your choice. Of course, you can also distribute your audio newsletter via our Playlist Player. This can be embedded into relevant web pages or shared via URL — perfect for email. Create a custom playlist today With custom playlists, it's even easier to get audiences in front of audio they're interested in and keep them engaged for longer. These three ideas should give you a good starting point, but we'd love to hear what else you come up with! Already signed u...
Demand for audio content is higher than ever, but what motivates people to listen? What causes someone to click play on an audio article, download a podcast, or opt for an audiobook? Understanding this is key to engaging and segmenting your audience effectively. Keep listening for the 5 key reasons why people listen to audio content. reason 1. to multitask In the US, the ability to multitask is the most popular reason for listening to spoken-word audio. 71% of monthly listeners say it plays a role in their decision. While reading pretty much requires our undivided attention, listening gives us more freedom. We can consume audio while we're getting ready, cooking, cleaning, exercising, or driving, for example. With so many things competing for your audience's attention, it's never been more important to create content that fits into a busy lifestyle, such as audio newsletters, narrated articles, and podcasts. Peter Charlton, NOVA Entertainment CEO, said: We really are living in the glory days for audio. Not since the advent of the Walkman in the early ’80s have we seen the same kind of exponential increase in personal audio consumption. We’re seeing considerable, ongoing growth in both audience size and time spent, which is a claim very few — if any — other mediums can make right now. However, it's not just about increasing productivity. 25% of podcast listeners listen to fill empty time — to keep their minds occupied and entertained. This means they'll consume content they'd never even consider in another format. reason 2. to concentrate Some people find it easier to concentrate on reading when they can listen along at the same time. This is known as immersion reading or audio-assisted reading. This bi modal learning technique is popular among language-learners, as it allows them to pair spellings and pronunciations as they go. But it can even improve comprehension and engagement in native readers, especially when it comes to complex texts. When you focus your sight and sound on the same thing, there are more ways to absorb the words — and less room for distraction. Book reviewer Arvyn Cerézo said: I decided to fire up my Kindle Paperwhite to read along with the narrator. Guess what, it was a eureka moment for me. After weeks of doing this, I think it facilitated my reading comprehension and made me understand the story better. Others find it easier to concentrate on audio alone. 60% of listeners say spoken-word audio allows them to process information more efficiently. Plus, studies suggest that audio can resonate more deeply than other content formats. UCL found that listening to audiobooks elicits a more intense physiological and emotional reaction than watching films or television, while the National Literacy Trust says that listening can aid comprehension. reason 3. they prefer listening to reading 56% of monthly spoken-word audio listeners say they prefer listening to reading. Even in 2018, when audio consumption was lower¹, 19% of Americans said they'd rather listen to the news than watch or read it. We all process information in different ways, and approximately 30% of Americans are auditory learners. These educational needs no doubt have a bearing on media preferences, especially when it comes to informational content such as news and reports. Emma Rodero, a communications professor at the Pompeu Fabra University in Barcelona, said: Audio is one of the most intimate forms of media because you are constantly building your own images of the story in your mind. Perhaps audio preferences stem from the oral tradition. Throughout human history, people have shared stories and information predominantly through speech — not writing. And even if you prefer reading or watching most of the time, you'll favor audio on some occasions — when your eyes are feeling tired, for example. reason 4. reading or vision difficulties Making your content available in audio is an effective way to improve accessibility, as it can cater to peop...
Learn how to elevate your text-to-speech audio. Whether you're creating audio from plain text or editing automated audio, you can use these tricks to make your content more engaging. 1. Mix and match your favorite voices You can change the voice associated with any title or paragraph in just two clicks. With a variety of virtual speakers at your fingertips, you can craft richer and more engaging audio using multiple voices. This feature is ideal for quotations. Not only can you give speakers more befitting virtual voices, but you can provide clear audio cues as to where speech begins and ends. You could also try using a different voice for your titles and subheadings. Plus, you can change your default voices in seconds. This means you can easily switch things up before creating your next audio. Our library includes over 550 AI voices across more than 130 language locales. We also offer voice cloning, so you can create custom voices that perfectly suit your needs. 2. Insert your own audio clips You can insert audio files into your text-to-speech audio without using an audio editor. Simply enter your text like usual then upload audio wherever you’d like it to play, and we’ll stitch it seamlessly into the final audio file. Combining synthetic spoken-word audio with other listenable content gives you the best of both worlds. Make your reports more trustworthy by inserting real interview snippets, provide clips of referenced audio to improve comprehension, or use sound effects and music to create atmosphere. When creating longer audios, it's a good idea to insert short transitional sounds, known as bumpers and stingers, to improve flow. Just like this one.
You've added audio narration to your written content. But have you updated your distribution strategy to match? Many publishers keep promoting their content the old way: like it's read-only. When sharing via social media posts, email newsletters, and other channels, they don't mention that there's an audio option. And so audiences without the desire or ability to read continue to scroll on by. If you advertise audio functionality up front rather than leaving visitors to discover it for themselves, you're more likely to attract clicks from those with listening needs and preferences — people who are on the go or too tired to read, for example. This can make a major difference to your engagement metrics, both offsite and onsite. Keep reading or listening for three distribution tricks that will boost traffic to narrated content. 1. Switch up your CTAs The easiest place to start is with your calls to action, or CTAs. Don't just tell people to read or learn more: tell them to go listen to your content. This simple trick can improve click-through when you're directing people to your content from: the page's meta description, which appears on search engine results pages; the page's open graph description, which appears on link previews; posts and comments on social media platforms or online communities; other pieces of content, such as related videos or guest posts; landing pages, such as content hubs or categories; and newsletters and other communications. You could also try incorporating separate CTAs. For example, in some of their newsletters, Twipe includes a 'Read more' button as well as a 'Listen to our article here' link. This allows them to measure initial interest in the audio version through link tracking and capture clicks they wouldn't have otherwise. 2. Provide listening times Providing reading time estimates up front can improve engagement with written content. As Maria Konnikova wrote in The New Yorker, quote, "The more we know about something — including precisely how much time it will consume — the greater the chance we will commit to it." End quote. When your content is listenable, you'll want to include the audio playback time too. Listeners are often busy multitaskers, so this information can be even more valuable in helping them manage their time. 3. Share audio snippets Sharing an extract lets your writing speak for itself. It gives readers a good idea of what to expect and helps convince them to click through. The same goes for audio content. Try creating short video trailers for your audio articles, which you can share on social media. This works particularly well on platforms with autoplay functionality, like Twitter and Facebook, because the sound will help grab users' attention before they scroll past. Start attracting listeners, not just readers When distributing narrated content, audio functionality is a USP that you should use to your advantage. Promoting it actively — rather than leaving site visitors to discover it for themselves — will help boost traffic and engagement. Looking for new ways to convert your content into audio? With our AI voices and embeddable players, providing quality narration has never been easier. Sign up free and see for yourself at beyondwords.io.
With demand for spoken-word audio on the up, 80% of media leaders are investing more in digital audio this year. If you're making the move into spoken-word audio, one of the main things you'll need to consider is human vs AI audio. While publishers like Zetland, The Economist, and Harvard Business Review have seen success with human voice over, publishers including Berlingske, The Japan Times, and Media24 are engaging audiences with synthetic speech. Some, like The Washington Post, use a mixture of both. Human voices can be more engaging, but this comes at a huge cost — one that's often unviable. With the quality of synthetic speech catching up to, and in some ways surpassing, human voice over, many find that the balance has tipped into AI's favor. Especially when it comes to audio articles and newsletters. In this article, I'm going to compare human vs text-to-speech audio production in terms of quality, cost, and time, to help you make an informed decision. Quality Human-read audio is generally considered more personable and engaging than AI audio, because people can add more emphasis and emotion into their speech. They can also manually check pronunciations and make thoughtful decisions on delivery. However, if you haven't had training in narration or voice acting, delivering clear and engaging speech yourself can be difficult. The sound can also be compromised by your recording environment and equipment. Hiring a voice actor and professional recording studio will give the highest-quality results, but this can be time-consuming and expensive. You may also have issues with achieving a consistent brand voice, because you will be relying on the availability of the voice actor. Another drawback with human-read audio is a lack of flexibility. Switching between multiple languages or voices means hiring and managing multiple speakers. This compromises your ability to choose the best voice for each piece of content you're producing. It's also impractical to edit human red audio after publishing. AI audio offers more consistency and reliability, as well as flexibility. With BeyondWords, you can easily update what's being said and switch between more than 500 voices across 130 plus language locales. There's even the option to create custom voices. This means you can clone your own voice, or the voice of a person on your team, to give audio a more personal touch. Or, you can work with a voice actor to create a unique and engaging brand voice. While AI audio is not as personable and emotive as human speech, progress is being made. And in certain cases, it's hard to tell the difference. The quality of your text-to-speech audio will depend largely on the AI voice itself. As well as having the option to create a custom voice, our users get access to an AI voice library featuring voices from Amazon Polly, Microsoft Azure, and Google Cloud. Subscribers can also use premium voices like Joe, which are ethically created in collaboration with voice actors. But the voice isn't the only thing that matters. AI voices sound better on BeyondWords because we use natural language processing algorithms to convert your text into speech synthesis markup language. This reduces the risk of pronunciation errors and allows for custom text-to-speech rules. Cost Human-read audio is traditionally expensive to produce. Of course, the cost will vary significantly depending on scale, how much work you do yourself, and how much you want to invest in quality. Technically, you can use your own voice, your phone or computer, and free software to create audio for nothing. You'll just need to account for your time and perhaps pick up some new skills. However, podcast producer Jeff Large explains that individual podcasters can spend over $2,000 on equipment, $799 on software, and $99 on hosting alone. Businesses that want to hire a podcast production team are looking at $1,000 to $15,000 per episode. To give some further context, if you are hiring a voice actor to rea...
Choosing the right AI voice is key to maximizing audio engagement. You need a lifelike voice that listeners enjoy listening to and find easy to understand. It should also reflect the nature of your content and brand. We've brought together 7 tips for choosing an AI voice that truly resonates with your target audience, so you can create automated audio that keeps listeners listening. 1. Choose the right language and locale The most important thing is that the AI voice speaks your language. But if your language is spoken in multiple regions, you might also need to consider locale. Our voice library covers 70 plus languages across more than 140 locales. For example, you can have your English-language content read in an Australian, British, Canadian, Hong Kong, Indian, Irish, Kenyan, New Zealand, Nigerian, Filipino, Singaporean, South African, Tanzanian, US, or Welsh accent. We also use different natural language processing algorithms for each locale. This allows us to ensure that voices pronounce words and other elements like a native would. 2. Think about representation Biases concerning accent, age, and gender can affect people's perceptions of voices. Listeners make judgements about traits like intelligence and trustworthiness based on accent alone. These often reflect stereotypes that exist in the general public. For example, one study found that female voice overs are considered more soothing, whereas male voice overs are considered more forceful. Rather than appealing to these biases, you can make your audio more authentic, engaging, and inclusive by focusing on representation. Millennials and Gen Zs in the UK agree that, as a culture, we’re more open to hearing from diverse voices than ever before. If you are converting authored content into audio, using the author's voice adds a personal touch. With our voice cloning service, you can create an AI voice that sounds just like you or a writer on your team. The next-best option is to choose an AI voice that best matches the accent, age, and gender of the author. Alternatively, represent the audience. People can better identify with voices that sound like their own, and tend to find them more trustworthy. This is particularly important when you're covering in-group topics, such as women's lifestyle or regional news. Our AI voice library includes adult male and female voices from 140 plus locales. We also offer a number of children's AI voices, which are ideal for making child-friendly audio. 3. Create demos Listening to voice demos can help you make a subjective call on how natural the AI voice sounds and whether it will suit your content. When you're choosing AI voices in BeyondWords, you can press the play button alongside to hear a short preview. The voice will typically say "Hello, you are listening to a preview of this voice", or a translation. You can also listen to these previews via our voice demo tool on beyondwords.io forward-slash voices. After shortlisting your favorite voices, you might want to create extended demos in our Text-to-Speech Editor. You can enter any text and set a different voice for each paragraph then review the audio output via our web player. This makes it easy to experiment and compare different options. 4. Listen to audio in your niche If you don't know what kind of voice will suit your content, it's helpful to listen to other audio in your niche. Whether it's delivered by a human or AI, this will give you a good idea of what voice types work. News broadcasting, for example, is associated with a distinct speaking style. This helps reporters to be clear, neutral, and trustworthy. As a result, similar voices are being used for AI narrated news articles. Amazon Polly has even developed specific Newscaster voices, which are available through our library. 5. Support voice actors Most AI voice providers don't pay commission to the voice actors who contribute to their voice models. We believe that voice actors should be fairly compensated for the l...
Your readers want to stay in the loop. But do they have the time? The battle for audience attention is fiercer than ever. We have to create content that fits into a busy lifestyle. That's where an audio newsletter comes in. What is an audio newsletter, and why create one? An audio newsletter is a listenable roundup of developments in a niche. It can be an audio version of a written newsletter or a standalone piece of content. Newsletters are typically delivered on a daily, weekly, or monthly basis, distilling a body of relevant updates into a manageable size. An audio format makes them even less time-intensive. Subscribers can listen while they're driving, exercising, or doing chores. All of this means that it's quick and easy for your audience to stay updated and engaged with your brand. This can help you get new customers or users — and keep existing ones. This is particularly important for media businesses operating on a subscriber-based model. 13 percent of news subscription cancellations result from subscribers having too much to read in too little time. Exclusive audio newsletters can help combat this unread guilt factor. Podcast charts demonstrate the demand for snackable audio. Daily news podcasts make up just 1% of podcasts produced, but account for 10% of top episodes. 44% of these podcasts take the format of a news roundup or micro bulletin. Adam Pasick, from The New York Times, said, quote: Newsletters are a great way to grow a subscription business. Subscribing to a newsletter is a healthy habit that can bear fruits down the line. End quote. How do I create an audio newsletter? If you already write a newsletter, you can directly convert future editions into audio newsletters. This is the quickest and easiest way to get started. However, it may be worth tweaking your text to create a more effective audio script. If you don't write a newsletter, or you'd prefer to build your audio newsletters from scratch, you'll need to put together a content strategy as well as individual scripts. Consider questions like: How often should I publish my audio newsletter? How long should each audio newsletter be? What do my listeners want to hear about? How can I keep my listeners engaged? What action do I want my listeners to take? You can record the audio yourself or hire a voice actor. However, this can be time-consuming and expensive. With our text-to-speech platform, you can produce engaging audio in minutes. Just choose your favorite AI voices, paste your text into our Editor, then hit 'Process audio'. If you send your email newsletters with Ghost, you can create the audio versions automatically using our Ghost integration. It's then easy to distribute, monetize, and analyze your audio. Where should I distribute my audio newsletter? Distribution is about getting your audio newsletter in front of the right people, in the right way, at the right time. The best approach will depend on your target audience and your wider marketing strategy, but these three channels are well worth considering. 1. Podcast feed One of the best ways to distribute an audio newsletter is via podcast feed. This allows you to publish on directories like Apple Podcasts, Spotify, and Google Podcasts, which reach approximately 485 million users worldwide. These platforms make it easy for users to find, listen to and subscribe to your audio newsletters. You can set up a podcast feed on BeyondWords in just a few minutes. Make sure to optimize your RSS tags to maximize reach and engagement. 2. Email It's not possible to embed audio files into an email, meaning that recipients can't listen within their inbox. However, you can use email to notify subscribers about your latest audio newsletter and direct them to the right place. For advice on sharing audio via email, check out our article 'how and why to make newsletters listenable'. 3. Website If you have a website, don't miss the opportunity to advertise your audio newsletter. Our embeddable playlists allow vi...
The tags in your podcast RSS feed provide key information about your show and episodes, affecting how and when they are displayed in podcast directories. Optimizing this meta data can therefore improve podcast visibility and engagement on directories like Apple Podcasts, Google Podcasts, and Spotify. It will not only make it easier for users to find your show and episodes, but make them more likely to listen. And, when you host your podcast on BeyondWords, updating your RSS feed is simple. Keep listening to learn about the tags you’ll need and how to optimize them. Show tags Show tags provide information about your podcast series. This means that they apply to the single channel element within your podcast feed. Title The title tag provides the name of your show, and is required on Apple Podcasts and Google Podcasts. This appears whenever your show or an episode is featured, so it’s important to choose something engaging and descriptive. Your title will affect which search results your show appears in, so it’s wise to consider what would-be listeners might search for. However, if you include lots of keywords purely for search optimization purposes, your podcast may be removed from directories, so tread carefully. You can also make it easier for listeners to find your show by naming it something unique, memorable, and that’s unambiguously pronounced and spelled. For example, numbers can create confusion, as listeners won’t know whether to search with digits or words. While the limit is 255 characters, it’s a good idea to keep your podcast title concise. This will make it more memorable and reduce the risk of it being clipped. The most popular length for podcast show titles is 16 characters, and 75% are 29 characters or shorter. When creating a podcast feed in BeyondWords, you can simply enter your show title into the corresponding field. Language The language spoken on your show. This is submitted as the two-letter language code defined by ISO 6 3 9 1. For example, English is E N. If you’re optimizing your podcast feed in BeyondWords, you can simply select the language from a dropdown list. Description The show description, also known as the podcast summary or blurb. This should explain what your show is about and what it has to offer, encouraging browsers to become listeners. You can include up to 4,000 characters. This includes any code, if you decide to submit HTML in a C DATA section. However, the average description is just 243 characters long. Podcast directories will typically display just a line or two, hiding the rest in a ‘read more’ section, so it’s important to front-load the most important details. Some podcast search engines take show descriptions into account, so it’s worth incorporating keywords and phrases where you can do so naturally. When creating a podcast feed in BeyondWords, you can simply enter your show description into the corresponding field. Link The main website or webpage URL associated with your podcast. This is optional, but it’s worth providing if you have one so that users can learn more about your show. When creating a podcast RSS feed in BeyondWords, you can simply paste your podcast URL into the corresponding field. itunes author The name of the person or organization authoring content for the podcast. In BeyondWords, you can easily add and delete multiple author names. itunes owner This show tag contains the email address of the podcast owner in a nested itunes email tag and the name of the owner in a nested itunes name tag. These details are used for account verification and admnistration purposes and are not displayed on podcast directories. If you’re creating your podcast feed in BeyondWords, you can simply enter the name and email address into the corresponding fields. itunes category The category or categories that your show falls into. This tag can help your podcast to appear in relevant categories, charts, personalized recommendations, editorially curated collections, and search r...
You can now customize the text on the BeyondWords Small Player, no coding required. This means that you are no longer restricted to using 'Listen to this article' — you can use the most effective call-to-action when embedding any audio online. There are lots of different ways you can use this feature to your advantage. Keep listening for inspiration, or head to the distribution section of your project dashboard to get started. 1. Break your audio down into sections It's good practice to break your writing down into sections, to make it easier for readers to find and digest the information they need — especially when it comes to long-form content. Same goes for audio. BeyondWords audio players already have a progress bar that allows seeking — in other words, users can move to different positions in the audio. However, you can make things easier by breaking your audio down into separate chapters. Step 1. Customize the Small Player using text such as 'Listen to this section' Step 2. Create the audio for a section using our Text-to-Speech Editor Step 3. Get the Small Player embed code for the audio Step 4. Paste the code into the corresponding section of your page Step 5. Repeat steps 2 to 4 for each section of content Publishing audio by section can also make it easier for users to read and listen simultaneously. 2. Insert audio clips for reference Whenever you're writing about something audible, it's helpful to provide reference audio. Let's say you're trying to clarify a pronunciation. You can create the audio in our Text-to-Speech Editor and then embed with custom player text like 'Listen to the pronunciation'. If you need the pronunciation of a foreign word, you can create another project under a different language. With a paid BeyondWords account, you can even embed other audio files, such as interview recordings or music clips. Step 1. Edit the player text to something relevant, like 'Listen here' or 'Listen to the clip' Step 2. Make sure you have the rights to use the audio, and edit if necessary Step 3. Upload the audio to your project in mp3 format Step 4. Get the Small Player embed code for the audio Step 5. Paste the code into the corresponding section of your page If you'd like to insert the clip within your audio version as well as the text version of your page, simply select 'Upload audio' when drafting or editing content in the Text-to-Speech Editor. 3. Improve engagement with other content types Now that you can change the default call-to-action on the Small Player from 'Listen to this article', you can improve engagement with other content types. For example, if you're creating audio versions of your landing pages, you can use a call-to-action like 'Listen to this page' — just like we have on the BeyondWords homepage. Try experimenting with different options and seeing the results in Analytics. The Small Player is used for automatically embedded audio, although Enterprise users can set a custom user interface using our JavaScript Player SDK. You can also use the Small Player when embedding audio manually, although you may prefer to use the Medium Player. This features the audio title as opposed to a customizable call-to-action. It also includes the project title and an optional image. BeyondWords audio players make it easy to maximize reach and engagement. There are also many other ways to distribute your audio. You can also share your audio via URL, create a podcast feed, download audio, and create a 'Latest audio' playlist. Visit beyondwords.io to create your free account and get started, or upgrade your plan to access advanced features.
6 December 2021. We've launched an exclusive AI voice named Joe. It's the new voice of BeyondWords, and it's available in beta to all our paid users! Joe was created in collaboration with Joe Coen, a British English voice artist, as part of our mission to make synthetic speech work harder for publishers, listeners, and voice artists alike. It's the AI voice you're listening to right now! The mission We want to give publishers more control over what they sound like, to help them better connect with listeners. That's why we've invested in voice cloning research and development. As well as enabling users to create custom voices, we want to bring unique voices to our voice library. Keen to adopt a new brand voice as part of our re-brand, we decided to work hand-in-hand with a British voice artist to create the first-of-its-kind AI voice service agreement. Our team wants to ensure that the voice artists we work with maintain control over their voice clones and are fairly compensated for their usage. So, we reached out to the Open Voice Network, a non-profit dedicated to voice industry ethics, to discuss the fairest approach to voice licensing. The human voice We found Joe Coen through the London Voice Boutique. With a soft, neutral intonation that would work well across content types, plus experience doing voice work for audiobooks, radio, and T V, we knew he would be the perfect fit for this project. Coen said, quote. I had heard some stories about how text-to-speech voices were taking work from voiceover artists. However, when the team approached me, they pitched me the idea of developing a voice with a contract that protected my voice I P. I was interested in pushing forward protection for voiceover artists in this industry, so I decided to go ahead. We recorded eight hours of audio over four days in a studio in North London. I had to find the right balance between expressiveness and consistency, to ensure that the resulting voice model would be versatile and engaging. End quote. The recordings were then passed to our text-to-speech research team for pre-processing. This involved removing errors and silences, converting to the appropriate format, and computationally extracting the thousands of characteristics that make up Joe's voice. This speech data was then used to train a voice model, which went through objective and subjective tests before being deployed. The AI voice Although we expect to make some improvements over time, we’re impressed with the outcome. We're proud to have broken ground by creating voice in partnership with a voice artist and hope other publishers find the voice to be a great fit for their content. Coen said, quote. It was amazing to hear the similarities with the AI voice compared to my natural reading voice! I'm looking forward to seeing how it's used — there's a lot of demand for professional voiceover, but recording isn't always feasible when there is a fast turnaround needed for daily publications. I'm excited to see how ethically made voice clones like these can support the careers of voice actors and make synthetic audio more engaging in these specific areas. End quote. How to use Joe's AI voice The Joe voice is currently available in beta. If you're a paid user who'd like to try it out, drop us a message on hello@beyondwords.io or via the chat feature, and we'll make the voice available on your account. Not a paid user yet? Sign up now on beyondwords.io or upgrade your account via your dashboard. You can then get access to Joe's AI voice and a host of other premium features.
30 November 2021. We’re so excited to share what we’ve been working on — and what we have in store. The major scoop? BeyondWords is the new name for a new mission. Plus, free plans! Why we rebranded For the best part of a decade, the team behind SpeechKit has worked to integrate scaleable audio into digital publishing. We made it our mission to bring audio to every article and we developed a toolkit for just that. Over the past 18 months, adoption of text-to-speech has skyrocketed. We’ve helped teams at News24, The Japan Times, and Fox News, among hundreds of others, to scale audio across their websites, newsletters, and apps. Audio players powered by BeyondWords now appear on 250,000 articles per month. We’re now ready to launch a new mission. A mission to blur the lines between reading and listening by creating frictionless AI audio experiences. Our hope is that the name ‘BeyondWords’ captures imaginations and the potential of what’s to come. Furthermore, our work with voice over artists has broadened our understanding of what needs to be done to make voice technology inclusive, safe, and transparent. We’re moved beyond our original remit of text-to-speech tools into breaking new ground in voice AI. New pricing plans We want to make the benefits of text-to-speech accessible to everyone, so we’ve rebuilt our pricing plans with every kind of publisher in mind. We now offer a Free plan, which allows you to create up to 30,000 characters of audio every month. That’s approximately four articles, nine newsletters, or one white paper. You can create audio automatically or manually, embed audio automatically or manually, create a podcast feed, and much more! With our Creator and Pro plans, you get access to more features and more characters. The plans are flexible too — you can add character boosters to increase your limit further. As always, our Enterprise plan offers large volume publishers the most robust text-to-speech service on the market. Advanced features for integrating and scaling audio include web and mobile player SDKs, programmatic audio advertising integrations with VAST, and dedicated technical support and account managers. What else is new? We’ve reimagined the dashboard. This puts all our audio production, distribution, analytics, and monetization tools, as well as our AI voice library, at your fingertips. If you aren’t already a member, why not create a free account and explore for yourself? Let us know what you think The team and I would love to hear what you think about our rebrand. We’d also be happy to answer any questions you might have. Just send an email over to hello@beyondwords.io or DM us on Twitter at @BeyondWords_io. Want to start creating audio with BeyondWords? Create your free account with no credit card details required at beyondwords.io.
12 November 2021. Instagram has added a text-to-speech feature to Reels, allowing creators to convert video captions into audio. This is seemingly in an effort to keep up with TikTok, which launched text-to-speech in December last year. As we discussed in a previous article, AI voices on TikTok are well used, but have received a lot of negative attention. And Instagram could go down the same route. For starters, it’s only offering two voices, both of which have an American accent: Voice 1, a female voice, and Voice 2, a male voice. It also experiences similar issues with inaccurate pronunciations and robotic-sounding speech. These limitations are understandable in the context of these platforms, but audiences are increasingly exposed to advanced synthetic speech, so their expectations are high. BeyondWords, for example, offers a library of over 720 voices across 64 languages, and the ability to create a custom voice. We also use natural language processing and speech synthesis markup language to ensure more accurate text-to-speech. And we’ve powered over one billion listens for over 120 global publishers. So, I expect Instagram’s text-to-speech feature to get its fair share of criticism. But it’s an impressive feature that will no doubt assist creators and their audiences. How to use the text-to-speech feature on Instagram 1. Open Instagram and go to the Reels camera 2. Create your video then select Preview 3. Tap the button to add a text caption 4. Tap the text bubble twice then tap ‘Text-to-Speech’ 5. Select a voice then tap Done 6. Make any other edits then share your Reel With BeyondWords, you and your team can convert any text into quality audio. We also offer all the distribution, analytics, and monetization tools you need. Create your free account today at beyondwords.io.
one: The average US adult listens to 1 hour 34 minutes of digital audio a day two: Spoken-word audio accounts for 28% of digital listening three: 43% of the US population listens to spoken-word audio every day four: The most popular spoken-word audio topic in the US is news or information, with 56% of listeners having consumed this type of content five: 54% of over twelves in the US have listened to an audiobook six: 31% of people across 22 countries listened to a podcast last month seven: 50% of monthly podcast listeners in the US are aged 12 to 34 eight: 88% of US adults listen to the radio every week nine: When they play an audio article, website visitors stay over ten times longer and visit 19% more pages ten: Audio article listeners are 32% more likely to engage in multiple sessions than non-listeners For more statistics and links to sources, please see our blog post on beyondwords.io.
8 November 2021. Please note that this post was published before SpeechKit rebranded to BeyondWords. We’ve released version 3 of the SpeechKit text-to-speech plugin for WordPress, and it’s packed with improvements and new features that make it even easier to engage your audiences with audio. Keep reading or listening to learn what’s new. Existing user? If you have auto-updates switched on, you’ll already have access to these new features. Otherwise, you’ll need to update the plugin via your WordPress admin. 1. Create audio versions in bulk You can now create and publish audio for multiple posts at once. This means you can easily make old posts listenable, or add audio to posts you’d previously excluded. Simply open the relevant posts overview section in your WordPress admin; use the checkboxes on the left to select which posts you’d like to request audio for; select Generate audio from the Bulk actions dropdown; then select Apply. The selected posts will then be processed into audio — this typically takes less than a couple of minutes per post. You can check the status of audios via the SpeechKit sidebar or your SpeechKit dashboard. Once processed, your audios will be automatically embedded into the corresponding posts via the Small Player, which you can customize via your SpeechKit dashboard. 2. Check audio responses at a glance We have replaced the SpeechKit Status column in post overview pages with a SpeechKit column. This shows you the response from our API — in other words, whether audio has been requested; whether the Player has been disabled; and if there has been an error. You can check the status of a particular audio via the SpeechKit sidebar in your post, or via the Content section of your SpeechKit dashboard. You can also use the SpeechKit sidebar to determine whether or not audio is requested, and whether or not the Player is displayed. 3. Easily share debugging data If you’re having technical difficulties with a particular post, you can now send us the debugging data. Simply open the SpeechKit sidebar on the relevant post, go to ‘Inspect’ then select ‘Copy’, and paste the copied details into your message to our support team. Plus much more... We’ve also made a host of other updates to improve the speed, stability, and security of our WordPress plugin. Want to make your WordPress content listenable? Our text-to-speech plugin makes it effortless, so start your free SpeechKit trial at speechkit.io today.
Lots of text-to-speech service providers use AI voices from Amazon Polly, Microsoft Azure, and Google Cloud Platform. But these voices sound best when you use them in BeyondWords. Why? It’s thanks to our natural language processing algorithms. Setting voices up for success AI voices can interpret text in two formats: plain text, or speech synthesis markup language, which is known as SSML. SSML tags provide extra information to the AI voice, clarifying pronunciations and improving speech flow. Using SSML therefore ensures a higher-quality voice output. Some text-to-speech services only allow you to input plain text, meaning you can’t achieve higher-quality outputs through SSML. Others, like Amazon Polly, give you the option to manually insert SSML tags. But this is complex and time-consuming. Let’s say the voice is mispronouncing the name Joe Biden. Fixing this requires an understanding not only of SSML, but of the international phonetic alphabet — symbols that linguists use to represent speech sounds. This is not feasible for the majority of publishers. BeyondWords, on the other hand, adds the SSML tags for you. Whether it is imported from your website in HTML format or manually imported as plain text, your content is automatically converted into SSML before being processed by the AI voice. This is made possible by a layer of natural language processing, also known as NLP, algorithms, which can programmatically read, analyze, and interpret written language. How our NLP works Our NLP uses a combination of rule-based and neural network–based techniques. Our deep learning models are trained on large data sets, which allow them to learn how humans convert particular text elements into speech and how this differs depending on context. This is particularly useful for resolving ambiguities in text. For example, in the sentence ‘I read the book’, our NLP identifies, through contextual features, that the word r e a d is most likely being used in the past tense, and so should be pronounced like red as opposed to reed. It applies an SSML tag accordingly, ensuring the AI voice gives the correct output. Also consider non-standard elements, such as dates. Our system can determine whether to read a number as an cardinal or ordinal numeral — for example, as twenty nine or twenty ninth — based on the usage context and other features. Without this NLP and SSML layer, the AI voice may not predict which pronunciation of an element is correct and simply output a naive best guess. This is why many text-to-speech systems struggle with ambiguities. A customizable and evolving NLP Our team of computational linguists make iterative and domain-specific improvements to the NLP. This means that voice outputs evolve even when the AI voices behind them stay the same, and can adapt to the needs of each BeyondWords user. Take the abbreviation NLP for instance. In the context of this article, it refers to natural language processing, but it is also the airport code for Nelspruit Airport. In medicine, it can be shorthand for no light perception. With custom built text normalization rules, we can ensure that our system delivers the most relevant result for a particular publisher. You can even add your own text to speech rules. Text-to-speech providers without a post-optimization NLP layer have no way to efficiently extend or make domain-specific customizations to an existing voice. If they wish to correct conversion errors, they must re-train the voice itself — something that comes at great cost and cannot guarantee accuracy, especially when it comes to unusual and idiosyncratic text-to-speech conversions. Here are some more examples of what our NLP can do: Apply phoneme SSML tags to ensure correct pronunciation of novel or complex words Apply lang SSML tags to ensure foreign words are pronounced accurately Apply sub SSML tags to ensure symbols, abbreviations, and acronyms are pronounced properly Identify tweets embedded within blockquote HTML elements, then ...
31 August 2021. TikTok's text-to-speech feature lets creators add virtual voiceovers at the touch of a button. It makes videos more accessible and eliminates the need to read. So why do so many people hate on it? What makes TikTok voices so annoying? Well, to start with, it's pretty common for TikTok's text-to-speech to get pronunciations wrong. 'Drop the mike' comes out 'drop the mick'. It can't work out whether R E A D should be pronounced like 'red' or 'reed'. And it definitely can't handle 'quinoa'. The lack of voice options is a problem too. Creators currently have access to just one region-dependent voice. This can create a jarring effect for the listener, where the voice doesn't match up with the message. Take, for example, Scottish TikTok. As on Scottish Twitter, lots of captions are written in regional dialect. They simply wouldn't work with the British-English text-to-speech voice. The North American text-to-speech voice has a Valley Girl speech pattern, which many consider to be annoying. Not only is Valspeak highly stigmatized in the US, but its chirpy nature simply feels inappropriate with a lot of TikTok content. And it doesn't help that this replaced a more liked voice — one that was removed because the voice artist said she never gave permission for her voice clone to be used. TikTok's synthetic voices sound robotic too. Even if you ignore mispronunciations, their tone and inflections often sound very unnatural. All of this is especially problematic when you consider that TikTok is a social media network. It's personal by nature. When creators use a synthetic voice rather than their own, it reduces the human connection. Plus, the read-aloud feature offers limited benefits to TikTok viewers. Only small amounts of text appear on videos and the platform is highly visual anyway, so it isn't empowering users to escape their screen or multitask. Nor is it having a huge impact on accessibility. Users can't really opt out of listening either, as text-to-speech is played automatically if enabled by the creator. Is this the best that text-to-speech has to offer? In a word, no. TikTok's text-to-speech feature is no doubt impressive. And with 1.4 billion views on #TextToSpeech alone, it's almost certainly raised awareness of the technology's benefits. It's updated perceptions too — for many, synthetic speech still brings Stephen Hawking's voice to mind. But it doesn't showcase the best of what modern text-to-speech has to offer. Which is fair enough: it's not a major feature, and the current version is fit for purpose. The most advanced text-to-speech can interpret and read text like a human. Often better. And while that might not be necessary on a social media network, it's crucial when converting articles, guides, and other written content into audio. BeyondWords text-to-speech At BeyondWords, we use natural language processing to apply speech synthesis markup language to text inputs. This acts like a virtual voice director, telling the AI voice how to pronounce elements like 'mike', 'quinoa', and 'read' properly. We also use advanced synthetic voices, developed using deep learning algorithms. These AI voices are trained on real human speech, ensuring more naturalistic inflection and intonation. We even offer voice cloning, allowing creators to "speak" in a voice that truly resonates with their audience. Our text-to-speech platform is specially developed for publishers. Want to try it for yourself? Create your free account on beyondwords.io.
Adding voiceover to your Substack posts can help you gain and retain subscribers — and you don’t even have to use your own voice. Why add voiceover to my Substack posts? Adding voiceover to your Substack posts can help you reduce unsubscribes by making it easier for audiences to engage. When subscribers can’t find the time to read your newsletter, they experience what The Economist called unread guilt factor: they watch unread emails pile up, can no longer justify their subscription, and are more likely to cancel. Northwestern University research backs this theory, citing frequency of consumption as the biggest predictor of subscriber retention in digital news. When you provide a listenable version of your newsletter, subscribers are more likely to engage. Audio is generally easier to fit into a busy lifestyle than reading — 71% of Americans listening to more spoken-word audio say it’s at least partly because it allows them to multitask. It also provides an alternative when subscribers simply aren’t in the mood to read. Narration can also help you grow your subscriber list. Making your Substack posts listenable helps you tap into the huge demand for audio content, allowing you to attract and engage people with listening needs and preferences. Plus, our research suggests that listeners are more valuable than readers. Creating audio versions of your Substack posts also gives you access to new distribution channels, most notably podcast platforms. This can help you build your following via platforms like Apple Podcasts, Spotify, and Google Podcasts. How to create voiceover for Substack posts There are three main ways to create audio narration for your Substack posts and newsletter: record yourself, hire a voice artist, or use text-to-speech. Option one. Record the audio yourself. Narrating your Substack posts yourself adds a personal touch, which can help you build deeper connections with your audience. It also gives you full control over vocal delivery, although you’ll be limited by your voice acting skills. You can create an off-the-cuff audio version in the time it takes to read your Substack post aloud, using nothing more than your phone or computer. However, if you want professional and engaging narration, you’ll need to invest in better production value. For high-quality sound, you’ll need good recording equipment and a suitable recording environment. It will then take you around 40 minutes of recording and 1 hour of editing to create a 20-minute solo podcast, depending on your narration and editing skills. As a result, producing your own audio can be complex, time-consuming, and expensive. And you may not be able to achieve the production value you desire. Option two. Hire a voice artist. Professional voice artists should deliver the highest-quality voice narration, helping you portray a professional image and maximize listener engagement. However, hiring, briefing, and collaborating with professionals can be time-consuming. Plus, unless you agree on a regular publishing schedule, you may find that turnaround times are relatively slow — a significant downside if you publish news content. Costs may also be high, although you can find well-rated and affordably priced voice artists on platforms like Fiver. Option three. Use text-to-speech. Using the BeyondWords Text-to-Speech Editor, you can convert Substack posts into audio in just a few minutes. The associated costs are relatively low, and you can select an AI voice that resonates with your target audience. You can even use voice cloning to create a custom voice. That way, you can speak directly to your listeners without putting the time and effort into manual recording. BeyondWords uses natural language processing and advanced AI voices to deliver the highest-quality text-to-speech, but you may find that human-read audio is more emotive and engaging. There are also service fees to consider, although text-to-speech is generally a cost-effective option. With a free accoun...
20 July 2021. Please note that this post was published before SpeechKit rebranded to BeyondWords. As consumer demand for audio increases, more publishers are turning to text-to-speech. AI audio publishing platforms like SpeechKit offer a quick and cost-effective way to make written content listenable. But how are visitors engaging with these AI audio–enabled articles? To find out, we’ve compared listener and non-listener engagement across more than 28 million sessions in which the SpeechKit Player was loaded last year. AI audio captures the attention of new users The average new user spends just 2 seconds on-site when they don’t engage with AI audio. This increases to 225 seconds (+11,150%) when they do press play. In an increasingly competitive online world, publishers need to do more to stand out and convince new readers they’re worth their time. Offering your written content in an audio format is an effective way to capture their attention. AI audio keeps visitors coming back for more Listeners are 32% more likely to engage in multiple sessions than non-listeners, suggesting that audio keeps users coming back for more. With a Northwestern University study showing that frequency of consumption is the biggest predictor of subscriber retention in digital news, AI audio can play a key role in sales and customer lifetime value. “The number of cases for how audio fits in with media companies’ subscription businesses is growing.” — Lucinda Southern, Media Editor at Adweek Plus, returning users are 38% more likely to press play than new users. This means audio articles are more popular with habitual visitors — visitors who drive the most revenue in both advertising- and subscription-based revenue models. When returning visitors do opt to listen, they stay on-site for longer (+688%) and visit more pages along the way (+11.48%). Listeners visit more pages The average non-listener visits 1.17 pages per session, whereas the average listener visits 1.39 pages (+19%) before leaving a site, according to our analysis. This finding suggests that listening experiences encourage visitors to stick around and explore. By providing content in a high-quality audio format, which some may find more accessible or engaging than text, publishers give themselves a better chance of keeping visitors on-site, potentially driving more ad revenue. Listeners spend longer on-site We found that non-listening sessions typically last just 30 seconds, whereas sessions involving audio last for 322 seconds (+973%). This means that users stay on-site over 10x longer when they press play. As well as boosting brand loyalty and revenue, this increased engagement could benefit your search engine optimisation (SEO) performance. According to Backlinko: “Google pays very close attention to ‘dwell time’: how long people spend on your page when coming from a Google search. The longer time spent, the better.” Making your written content listenable could therefore improve your search engine rankings, meaning that more people discover your content organically. Audio engages at every age Our analysis backs the established idea that younger demographics gravitate towards audio content. 18-to-34-year-olds represented 28.02% of our sample but 37.92% of listening sessions, making this the age group most likely to press play (1.5x more likely than 35 and overs). Audio engagement also had the biggest impact on session duration here (+1,109%). However, it was listeners aged 55+ who remained on-site for longest, clocking an average session duration of 407 seconds. This group also visited the most pages per session (1.765), which represented the biggest increase versus non-listeners (+48%). This suggests that engaging older people with audio can result in the biggest pay-off. Audio interaction also had a positive impact on the engagement of 35 to 54s, corresponding with a 2% increase in pages per session and 847% increase in session duration. SpeechKit is an AI audio publishing platfo...
In 1985, Stephen Hawking had a life-saving tracheostomy that took away his natural speaking voice. A.L.S., also known as Lou Gehrig’s disease or motor neurone disease, had already caused his speech to slur and affected his ability to move. He communicated by raising his eyebrows when someone pointed at the right letter on a spelling card. That changed when Walter and Ginger Woltosz, founders of Words Plus, donated a communication system called the Equalizer. The married couple had originally begun developing it for Ginger’s late mother, Lucille Evans, who had ALS. The computer program scrolled through common phrases on a screen, and Hawking could select what he wanted to communicate with the touch of a button. When he submitted a message, it was processed by a speech synthesizer called the Speech Plus Call Text 5010. Its male voice had an American accent. The accent of Dennis H. Klatt. The man behind the voice Massachusetts Institute of Technology researcher Dennis Klatt had been working on speech synthesis since the 1960s. He developed an algorithm called Klatt Talk or MI Talk. This had three voices — ‘Perfect Paul’, ‘Beautiful Betty’, and ‘Kit the Kid’ — created using hours of recordings from himself, his wife, and his daughter. They were first released in 1984, as part of the DECtalk speech synthesizer. ‘Perfect Paul’ would soon be used by the Speech Plus Call Text 5010 synthesizer, too. Joseph Perkell, a colleague of Dennis Klatt, told Witness History, quote. The first time I really understood that Stephen Hawking was going to be using Dennis Klatt’s speech synthesizer was when I heard him talk. I thought, wow, I’m watching Stephen Hawking and out comes Dennis’s voice. It was kind of startling. At the time, the quality of it was about as good as you could get. Compared to the other schemes that were being developed in other places, [Klatt’s] clearly sounded the best. End quote. While working on technology that would give Stephen Hawking a voice, Dennis Klatt was losing his own. Thyroid cancer affected his vocal cords, and he spoke with a hoarse and raspy voice in the last decade of his life, before losing the ability to speak altogether. He died in 1988. His voice lived on. Professor Hawking used the Speech Plus Call Text 5010 until his death in 2018, despite the fact that he had been offered upgrades. In fact, when he needed a new synthesizer — two decades after Speech Plus had gone out of business — his team went to great lengths to restore Perfect Paul. "I keep it because I have not heard a voice I like better and because I have identified with it,” Hawking said in 2006. While many would describe Stephen Hawking’s synthetic voice as robotic-sounding, that didn’t stop him from communicating the most complex of ideas, whether he was delivering lectures at Cambridge University, conducting television interviews, or making speeches at NASA. Nor did it hold him back from conversation. Plus, Hawking’s voice became one of the most famous and recognisable in the world. It’s little wonder that filmmakers working on his 2014 biopic, The Theory of Everything, wanted to get it right. Screenwriter Anthony McCarten told Variety, quote. We spent a lot of time and money trying to reproduce the voice, but we never got it. End quote. Fortunately, Hawking was so pleased with the pre screening that he gave permission for the filmmakers to use Perfect Paul, which was now trademarked. Eddie Redmayne, who played Hawking in the movie, said, quote. With his specific voice, it’s an actor’s dream. You’re one step closer to the truth. End quote. Synthetic voices today Text-to-speech has come a long way since Dennis Klatt developed Perfect Paul. Klatt engineered his speech synthesis algorithm manually. Based on the parameters of his own voice, he formed rules to shape computer-generated sounds into speech-like sounds. This technique was ground-breaking at the time. Now, we have computer algorithms that can learn complex voice models with millions...
30 June 2021. Please note that this post was published before SpeechKit rebranded to BeyondWords. This is the SpeechKit changelog for June 2021, where you can learn about the latest updates and improvements to our AI voice publishing platform. In addition to overall performance improvements this month, we introduced a few new features: Added webhooks You can now use webhooks to get notifications when an event happens. Webhooks are particularly useful for asynchronous events, like when an audio has been processed, updated, or deleted. For more information, see our webhooks doc. Added multi-voice asset support to the API The SpeechKit API now offers support for multiple AI voices within the body of a single audio asset. This means that you can switch between speakers within a particular piece of content — ideal for: Quotes, which can be read in a different (and more appropriate) voice so that they are distinguished from the main narrator’s Sections of foreign-language content, which can now be processed and read in a native accent Subheadings or other features that you wish to distinguish using audio For guidance using this feature, see the Multiple voices per audio section in our API introduction doc. Added Categories selection to the WordPress plugin settings WordPress plugin users now have the option to automatically enable or disable audio by post category. This makes it easier to ensure that audio is processed for posts that you want — and not those you don’t — so you can use your audio credits effectively. For example, you could prevent posts in a ‘Videos’ category from being processed into audio, while ensuring that all your ‘News’ posts are automatically made listenable. If you are an existing user who wants to enable this feature: Login to WordPress and go to Plugins. Find the SpeechKit – Text-to-Speech plugin. If you haven’t already installed Version 2.15.1 or later, click update now. Click Settings underneath the SpeechKit plugin. Find the new Categories section under Advanced Settings. Tick the categories you wish to automatically enable audio content for. Click Save Settings at the bottom of the page. If you wish to override the default setting for a particular post: Open the post you wish to edit. Click the ⋮ icon in the top-right, navigate to the Plugins section, then click SpeechKit. Tick or untick Display Player to override the default for this post. For more information, take a look at our SpeechKit for WordPress doc. Added isAudioReady(object) to the JavaScript player SDK It typically takes no more than a few minutes for SpeechKit to process your content into audio, but if you want to make sure that the JavaScript player doesn’t render until audio is available, you can do so using the isAudioReady(object) method. For more information, take a look at our JavaScript player SDK docs. If you have any questions or suggestions, get in touch on hello@speechkit.io. Not yet a Speechkit customer? Our platform makes it easy to turn your written content into audio content, and it’s constantly evolving. Start your free 14-day trial.
24 June 2021. Please note that this post was published before SpeechKit rebranded to BeyondWords. Last year, after hearing of the success of other businesses, we decided to make an application for an Innovate UK Smart Grant, to support the development of custom AI voices. We’re pleased to announce successful funding! What is the Innovate UK Smart Grant? Innovate UK, part of UK Research and Innovation (UKRI), drives economic growth by supporting research and development at companies that show potential for considerable economic impact. The Smart Grant program awards government backing to “game-changing”, disruptive ideas that can provide evidence for future profitability. Our funding goals Our application focused on the development of a voice cloning system that allows us to quickly and efficiently develop and deploy custom text-to-speech voices. These voices allow publishers of any size, from bloggers and institutions to global companies, to transform and enrich their text content using audio. Our mission at SpeechKit is to build the best text-to-speech tools for our customers, putting audio on every article, every story, every report...every readable piece of content on the internet. For this to happen, publishers and companies need to control what they sound like. Furthermore, custom voices, tuned to the ears of a particular audience, resonate more deeply. Additionally, for AI voices to become ubiquitous throughout online publishing, we need to make this technology affordable for customers large and small. Grant funding allows us to develop more efficient training models, requiring less data at a smaller cost. The application process For founders looking into Innovate UK funding, here are some insights into our experience. In general, we’ve been really impressed with the Innovate UK process. We were required to give more detail than we originally anticipated, but bureaucracy has been minimal so far. We used a consultant, Grantify, to help us draft the application. The initial application process took approximately 20 hours, spent getting the plan on paper (within the word limit) and doing the required market research, technical documentation, budgeting, risk analysis, and work package plan. This required the four project participants to work together, remotely, iterating the plan. Despite being tedious at times, it was a really worthwhile exercise, and not just for the purpose of winning funding: it helped us articulate a vision for AI voices we were yet to put on paper. The questions in the grant also helped us to build our plans for our seed raise, which also benefited from the Innovate UK stamp of approval. Screenshot of Innovate UK's response to our application, reads: This is an excellent application. The team have the capabilities and experience to deliver the project. The project plan is realistic and has been well thought out. The funding should represent excellent value to the tax payer. A UK Innovate judge's response to our application Success! Our application was marked successful on 5 February, subject to thorough financial checks which took another two to three months. We were appointed a Monitoring Officer, who we’re to report to every quarter for the duration of the 12-month project, and who reports to Innovate UK, confirming our progress as we submit it. Innovate UK reimburses 50% of costs related to the project as they’re reported. These are exclusively R&D-related: marketing, sales, and other technical costs are not covered. We’re only permitted to employ workers in the UK, or who are employed through PAYE. Our team are massively grateful to be supported by Innovate UK and the UK government. We’re excited to reinforce our business with the investment and push custom voices to the next level. Interested in developing a custom AI voice with SpeechKit? Our article-to-audio technology makes large-scale audio publishing simple and affordable, and we can develop custom AI voices that truly speak to your audien...
18 June 2021. Lots of us read about, discuss, or even watch Netflix series without ever knowing the intended pronunciation of their names. Remember when the company claimed that English speakers should say emi-lee in par-ee? One of the most common difficulties is with foreign titles. Netflix often renames these shows (The Killing is named Forbrydelsen, which literally translates to The Crime, in its native Denmark), but when they're left untranslated, it can be difficult to know if you're pronouncing them correctly. Here, we've used our text-to-speech technology to help you get to grips with the standard pronunciation of five popular Netflix titles. Lupin Netflix has just launched five new episodes of Lupin, a French mystery thriller that has been Number 1 in Netflix’s Top 10 in most countries across the world. It’s the most-watched non–English language Netflix Original, with 70 million views. The show’s protagonist is inspired by the adventures of Arsène Lupin, a fictional thief created by French writer Maurice Leblanc. Mispronunciation of the eponymous show’s name is so common, Netflix have launched a video to try to solve the issue. You might not be able to nail the French accent, but loo-pan is less of a faux pas than loo-pin or luh-pin. Borgen Borgen is a BAFTA award–winning Danish political drama named after the nickname given to Christiansborg Palace, the home of the Danish government. It’s pronounced more like the English word born as opposed to what it looks like: borg-en. You have to say it “like you have a hot potato in your throat”, Nordic Noir Tours guide Dieuwetje Visser told the BBC. Fauda Fauda is an Israeli TV series based on the creators’ experiences in the Israel Defence Forces. The show has a 100% critics score on Rotten Tomatoes. The name Fauda comes from the Arabic word fawḍā, meaning chaos. English speakers might think it's pronounced for-da, but it’s more like fow-da, almost rhyming with louder. Atelier Atelier is named Underwear in its native Japan, but this drama was retitled for English-speaking audiences. An atelier is a designer’s workshop, so both names are equally appropriate when you consider that the show is set in a lingerie design house. As the word atelier has French origins, we recommend that it’s pronounced a-telli-ay. Fallet Fallet is a spoof of the Nordic noir genre, in which Swedish and English detectives are paired up to investigate the murder of an English man in Sweden. The show stars Adam Godley, known for appearances in shows like Breaking Bad, Suits, and The Umbrella Academy. This Swedish title translates into The Case. You’d be forgiven for thinking Fallet rhymes with ballet, but fal-et is the way to go. BeyondWords text-to-speech technology detects foreign words so that they're pronounced correctly every time. To turn your written content into quality audio, create your free account on beyondwords.io.
4 February 2021. Please note that this post was published before SpeechKit rebranded to BeyondWords. The speech services market is young but already being disrupted. How do you compare company X or company Y in such a fast moving environment? In this post, we lay out how we see this rapidly evolving market and what differentiates one company from another. First, some background... Text-to-speech has often been overlooked as a revolutionary new media format. Over recent years, as with much nascent AI, text-to-speech has been scrutinised over error rates, poor voice quality, and its ‘robotic-ness’. After 4-years of working to improve this technology, listening through hours upon hours of audio data, our team is well versed in these imperfections. An early test trained on voice data from Desert Island Disks. A recent test trained on just 10-minutes of public data. However, great improvements have been made - by our team and by our peers. Last year, Nic Newman at the Reuters Institute predicted audio articles would become ‘standard’. New neural voices from SpeechKit, Amazon, and Google, among others, are coming closer to human speech than could have ever been imagined. Furthermore, text-to-speech is starting to show very real results. The average engagement time of a listener is more than four times that of a non-listener. Text-to-speech is being adopted to improve customer value and fast-track digital transformation. Below we’ve segmented the market into three categories to make for easier comparison. Cloud Service APIs Services, such as Amazon Polly, Google WaveNet and IBM Watson, provide text-to-speech APIs as part of their cloud service platforms. These products have been developed over the last 10-15 years, producing some of the most robust voices on the market. Early adopters, such as Bloomberg and The Globe and Mail, adopted Amazon Polly across their publishing to give their readers the option to listen. Whilst high-quality, these services lack utility and customization; they’ve produced very powerful APIs, yet haven’t created any tools (hosting, CMS integrations, analytics, monetization) to help publishers get the most from text-to-speech. Strengths: reliability, voice-quality, cost. Weaknesses: ease-of-integration, customization, publishing tools. Custom Voice Startups Speech startups, such as Resemble and Sonantic, aim to out-compete those cloud services mentioned above, by creating custom synthetic voices. Similarly, these companies have deep expertise in Machine Learning, developing high-quality voices for sectors such as gaming and customer service. Similarly to cloud services, whilst these companies have developed advanced APIs, they’ve neglected publishing tools. Neither of these examples provides audio distribution tools, analytics or monetization options for publishing at scale. Strengths: voice-quality, customization, ease-of-integration. Weaknesses: publishing tools, cost, scalability. Text-To-Speech Startups Startups, such as Trinity Audio and Play.ht, are our most direct competitors. These, like ourselves, are developing tools around text-to-speech, creating value for publishers and bloggers. Trinity Audio have focused on monetizing APIs created by those cloud services mentioned above. They’ve built ad insertion technology and partnered with multiple audio DSPs to improve fill-rates and potential CPMs. Their business model aims to provide revenue back to publishers who adopt their tech, while they take a cut. Whilst this is sound in principle, we believe a focus on voice quality, infrastructure and customer service most benefit our customer’s long term. Strengths: monetization, cost, scalability. Weaknesses: ease-of-integration, voice-quality, customization. SpeechKit In comparison, SpeechKit has focused on becoming the ‘full-stack’ service for automated audio publishing. We help news publishers, institutions, corporate businesses and bloggers publish their written content in audio, automatically and at sca...
23 October 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. As reported by the 2020 Spoken Audio Report, produced by NPR and Edison Research earlier this month, we are listening to more spoken audio than ever before. Spoken audio's share of listening has grown 30% over the past 6 years, 8% this last year. Text-to-speech technology has come a long way in the last 6 years. Gone are those days when interacting with automated voices was confined to navigation, service centres and buggy early AI-assistants. Text-to-speech has proliferated into 50+ languages, and many more different voices, tones and styles. SpeechKit works primarily with news publishers who, until now, have been limited to adopting off-the-shelf voices, provided by Amazon Polly and others. These voices have worked great to increase engagement with news articles and provided readers with an option to listen to audio articles. However, there’s room for improvement. Cape Town, home of News24's local voice. Image credit Zoë Reeve. Resonance keeps us engaged All brands, whether that be in news media, consumer or any other sector, want to talk to their audience in a voice that resonates. This is the sweet spot of accent, intonation and style which satisfies listeners. Its why great voice artists can charge what they do and (what we’re finding at SpeechKit) it’s what keeps listeners engaged with audio articles. Studies show that listening comprehension improves when a sentence is spoken in someone’s native accent. We’re more receptive to voices with meaning to us. Having lived much of my childhood in Australia, I live in the UK full-time now, yet my Siri is programmed to speak to me in an Australian accent. For one reason or another, it makes that experience personal – it gives it context and meaning. I find myself enjoying listening to the responses to my (less than meaningful) requests. Custom localized synthetic voices We’ve been helping publishers adopt audio articles for the past 3 years and we’re making good progress. Most recently we’ve started to develop custom synthetic voices for our customers who hope to create a voice that better resonates with their audience. Our first publisher, News24, South Africa’s largest news platform needed a South African-accented voice that could pronounce local names in Zulu, Xhosa and Afrikaans. We developed the voice using a new machine learning technique to model realistic voices, requiring less training data than conventional methods. The voice was launched as a premium feature on a new digital subscription product at the beginning of August. During that month we observed an 407% increase in audio engagement on News24.com when compared to publishers using conventional text-to-speech voices. Furthermore, over the past two months average listen length per article has increased to 2 minutes and 12 seconds. We’re confident that custom voices, tuned to the ear of the audience, are going to introduce new growth to text-to-speech audio articles. Advances in voice training are providing the efficiencies needed to allow for specialization and providing publishers the quality they need. Technology is allowing us to address the subtleties of voice, encouraging the further adoption of audio articles. More about SpeechKit At SpeechKit we’re helping 100’s of news publishers to automate audio versions of their news articles. Sign up for a free trial and instantly start engaging the audio generation.
1 October 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. For the past 5 years, SpeechKit has been developing text-to-speech tools for digital publishers. Our guiding principle has always been to improve text-to-speech integrations and to create audio for every story. We want more than anything to create seamless audio experiences — that make listening to the content we crave effortless. This week, we release a brand new minimal audio article player. Designed to fit inline with existing publishing and branding, the sleek new audio player sits atop each article. We’ve designed it with customizable features for enterprise use — such as play and pause buttons, progress bar and new background colors. We’ve also added multi-speed listening to allow for faster or slower listening speeds. Minimal design text-to-speech. The new audio article player comes as default on new projects and we’ll be transferring existing customers over the coming weeks. We’ve outlined a number of updated features below. For more information on SpeechKits integrations please get into touch with someone on our team. Player features We designed the new player taking feedback from customers, minimising integration and design hassles. Once integrated using one of the SpeechKit integrations tools — WordPress, Ghost, API and RSS — the player instantly serves audio alongside your content. CMS Integrations — The SpeechKit player integrates into your existing digital publishing tools for maximum efficiency. Once integrated, audio is published simultaneously alongside each piece of content. We’ve built integrations that fit the spectrum of digital publishing, including plugins for WordPress and Ghost, a SpeechKit API and RSS feed ingestion. Audio Analytics Tracking — Impressions, listens and duration are tracked on every audio article. Customers of SpeechKit benefit from real-time tracking available through their project’s dashboard. Track trends and performance of audio articles as your readers turn to listening. Alternatively, integrate with Google Analytics to track audio within your existing GA dashboard. Multi Playback Speed — We updated the player to include multi-speed playback. Speeds now available include x0.5, x1.0, x1.25, x1.5 and x2! Everyone’s got their style when it comes to listening. I’m a fan of the gentle and melodic x1.25! HLS Adaptive Streaming — A priority for us at Speechkit has always been to optimize load speed and minimize the latency of audio article players. The new player includes HLS Adaptive Streaming, meaning audio quality adapts depending on the listener's available network data. This feature addresses on-the-go users moving in-and-out of good network as they go about their days. JavaScript iFrame — Another speed-related design principal; reduce page loading speed. We never want our player to add to page loading, so we built the player in an iFrame which waits for the other elements on the page to load. Once it’s the turn of the iFrame the player loads in 110 milliseconds (as of last speed test), a 164% improvement on the existing player. More about SpeechKit At SpeechKit we’re helping 100’s of news publishers to automate audio versions of their news articles. Sign up for a free trial and instantly start engaging the audio generation.
7 September 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. Emerging Market Media is a New York-based digital media firm. Founded by Dawn Kissi, an award-winning journalist, and entrepreneur, the firm publishes news, analysis and opinion pieces on the world's emerging economies on its flagship website Emerging Market Views. Serving a global audience, Emerging Market Views chose to implement Speechkit's technology to offer its community a new option of digesting news. We recently sat down with Dawn to learn more about the platform, and how Speechkit has helped Emerging Market Media leverage its brand. Why did you choose to add audio articles to your publication? Today, it seems more and more our media diets are overflowing. Between newspapers, magazines, blogs, podcasts and of course social media feeds there is a massive deluge of information hitting us every day. Rather than add a totally new vertical at this point in time, we felt an easy to access and easy to use product would be a better fit for us right now. The audio option allows us to keep offering our coverage free of distractions while keeping our dedicated readership on the page, and in many cases, looking for more while engaging with a particular item. How does the behaviour of a listener differ to that of a reader? Our readership and audience are really global. And for the most part, they are busy working professionals within financial services--namely asset management, family offices, private banking, hedge funds and more. They have access to some of the most real-time, market-moving and in-depth data available today. Our position has always been to offer deep analysis, spotlight issues in some under-reported economies that are of interest to our communities and of course, relevant news with some editorial flair. Those reading, for the most part, know what to expect. Audio is a big compliment for both the reader and us, as publishers. Those that listen, tend to linger longer--the jump to the next article and once there again listen. Our longstanding, loyal readers at this point know what they are looking for, those running audio tend to be newer, and ultimately convert. How will you quantify the success of an audio strategy? We are small, and we are niche. That said, any offering that we present must be not only relevant, but worth it for both the reader and user, as well as us. Our investments in audio have paid off, far more than we had hoped or expected. Where and how do you see the audio format developing at EM Views? Audio will for sure remain part of our editorial offering, and there are plans to increase and enhance how we offer it. Today we are looking to implement audio articles within our newsletters and special editorial sections. For niche publishers, audio articles are a smart and simple option for diversifying an editorial product, and ultimately, a revenue-generating one. Visit Emerging Market Views to experience a rich, audio offering from this emerging global publisher.
6 August 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. SpeechKit offers a range of audio content management tools alongside cutting edge text-to-speech technology. Digital content publishers use SpeechKit to distribute podcasts and other pre-recorded content, taking advantage of hosting and distribution options. The SpeechKit Platform By harnessing the full capabilities of the SpeechKit platform, newsrooms, bloggers and agencies can manage all of their audio needs under one roof allowing for an easier and faster audio editorial process. This eliminates the need for multiple tools, platforms and dashboards. In this post we'll outline some key features. Players SpeechKit audio players are designed to fit inline with web and mobile content. Our players work for both automated publishing (through WordPress, RSS or API) and/or with Mp3s uploaded directly into the dashboard. To get started set up a new project in the SpeechKit dashboard. Choose Audio Inputs To set up a non-automated project - for podcasts or other content - create a new ‘plain-text’ project. Select a voice in case you choose to add text-to-speech articles too. More info on setting up new projects in the docs. Uploading Audio Once a project is set up audio can be added using text-to-speech or by uploading pre-generated audio, podcasts or Mp3s. Hosting audio content on SpeechKit lets you distribute it through audio players and through other distribution channels - such as Spotify, Apple Podcasts, and Google. Upload Mp3 In your new project, to upload a piece of audio content select ‘Add Article’, select ‘Upload Audio’, give your audio a name and hit ‘Upload Mp3 File’. Distribute From SpeechKit you can distribute audio to players embedded on your websites and apps, to Podcast Platforms and to Smart Speakers. Once you have audio you want to distribute head to the ‘Distribution’ tab in the dashboard. If you want to embed a single piece of audio or Mp3 you can do this from the content tab. Click on the ‘< >’ icon next to the article name to get an embed code for that piece of audio content, or use the URL to embed into Medium and other sites. A SpeechKit Playlist Player From the distribution tab you can set up playlist players, feeds to Spotify and Apple Podcasts, as well as to Alexa and Google Home. Select a distribution option and follow the prompts to get started. Check out this blog post for info on setting up feeds to Podcast Platforms. Analytics Lastly, using SpeechKit to publish audio will provide you access to data on how audio is performing. The analytics tab in the dashboard tracks listens, audio engagement, device engagement and more. If you are a customer of Google Analytics you might also want integrate SpeechKit reporting with your GA dashboard. To integrate audio into your Google Analytics access the docs here. The features of the SpeechKit’s audio content management system are designed to fit the needs of our publishing customers. If you have any question or feedback please contact support@speechkit.io.
29 June 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. Over the last few years, news of streaming wars between tech giants have distracted interest away from an interesting trend; in years since 2014, music streaming has decreased 5%, whilst spoken-word listening has increased by 20%. As a population, we’re turning off the tunes in favour of information and stories. Audio is being utilised by a hungry, connected and agile modern media consumer. Furthermore, new audio formats that fit modern life are meeting this insatiable desire for information. In the news industry, newsrooms around the world are ‘pivoting-to-audio’ both to engage new demographics and to improve engagement with existing ones. The recent high profile acquisitions of Audm by the New York Times, Gimlet and Anchor by Spotify, combined with the recent release of Audio articles to Apple News+ are clear markers of growing interest/momentum in spoken news audio. In this post, we want to outline the drivers of change, the valuable benefits to newsrooms and what’s next to come. Complete News Packages Over the last decade, we’ve observed a trend away news snacking, clickbait and tweet-sized news. As a recent Twipe Mobile study reveals, even in those younger demographics most attuned to snacking, an equal portion of the audience prefer to consume their news in complete packages. News readers who lend themselves to this behaviour tend to be lucrative sources of reader revenue. As a result, newsrooms and startups, such as Tortoise, are developing news products that reflect this trend back towards ‘slow news’. Spoken word lends itself to this not-so-new behaviour. Now that sitting at the breakfast table with a print edition no longer fits modern news habits, audio is filling the gaps. In routines interrupted by subway doors, slow pedestrians and intermittent instant messaging, spoken audio offers a refuge. As more text-based news is brought online (to audio), more personalisation and richer listening experiences are introduced, expect listening to take an even larger role. Perhaps that traditionally offered by radio. Subscriber Benefits We first heard evidence of audio’s subscriber benefits whilst listening to The Economist’s then Head of Strategic Product Development, Denise Law, describe the ‘unread guilt factor’. The Economist had recognised as early as 2007 that a primary driver of subscriber churn was the feeling of guilt brought on by the sight of a stack of unread editions. For their solution they turned to audio articles, narrated by paid voice actors. This has been an immensely successful strategy for their publication along with others, such as Zetland, where the audio edition is more popular than the written one. A strong ability to engage and retain newsreaders makes audio a valuable tool. From our own data, we find that audio articles are incredibly popular with return users — 48% of those who listen are returning. Another recent Twipe Mobile study found that format variety is the main factor leading to a decision to subscribe, even when compared to editorial coverage. As the market for reader revenue becomes more competitive, listening will become ubiquitous with paid news subscriptions. Valuable Demographics Digital audio presents a valuable opportunity to engage a younger audience. In the UK, younger age groups are four times more likely to listen to podcasts than over 55s, as reported by Reuters. Newsrooms are turning to audio to reach the next generation of paying newsreaders. Operations like the NYT’s ‘The Daily’ and the FT’s ‘News Briefing’ are being used to draw in younger news consumers on free platforms — building brand loyalty and securing future revenue. We observe a similar trend in audio articles. From our own data, 65% of audio article listeners are 44 and under (compared to 49% of newsreaders). Furthermore, those in the 18–24 bracket are 51% more likely to engage with audio, when compared to the ave...
6 March 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. A new term has appeared in digital publishing over the past 12-months. Audio Articles are just what their name describes; news articles in audio. Over the past few years, demand for news content delivered on-the-go has exploded. The number of people in the US listening to spoken word audio increased 20% in the last 5 years (during a time when music listening shrunk 5%). Up until recently, podcasts seemed to be the only valid answer to rising consumer demand for audio content. Yet podcasts deliver a set of unique challenges to newsrooms. How do you effectively replicate quality news coverage in podcast format? You may need to hire new talent to host a daily show, build a production studio, negotiate new distribution deals with the major podcasting platforms, find a suitable audio hosting provider and finally build an audience. All this for an audio format that proves financially lucrative for just a select few, due to the high level of competition in the Podcast market and platform oversaturation in the national and world news podcast segment. Audio News Articles provide a simple and effective alternative to conventional podcasts This format, proving to be popular with many publishers, takes existing articles and converts these into playable audio files that are served alongside their written stories. No need for new strategies, setup costs, or marketing. The average completion rate for an audio story is 90%. Hakon Mosbech - Zetland (Nieman) This new format provides readers with the option to listen to any article. Competition for time has never been higher – and tools to reach news on-the-go are increasingly valuable. Furthermore, readers are finding this useful. At SpeechKit, we’re seeing increases as high as 1200% in time spent on page when Audio Articles are implemented on news websites. Learn more here. Audio Articles are proving a useful tool for publishers searching for increased engagement with their stories. With SpeechKit, publishers of any size can implement natural sounding audio within minutes. To find out more, signup for a free 7-day trial, or get in touch with our team using the form below.
13 February 2020. Please note that this post was published before SpeechKit rebranded to BeyondWords. SEO is the favourite topic of internet marketers. We write about SEO in the hope of improving our SEO. Like website Inception — search term dreams within search term dreams — or a nightmare of post after post of rambling nonsense. However, search engine optimisation IS important! Search has become one of the most crucial entry points of news in recent years. The largest search engine, Google, has come under major scrutiny from publishers. Indexing on Google can determine the revenue of those sites that it either ranks higher or lower. Introducing Audio SEO Back in May 2019, at Google’s I/O conference, the company made a series of announcements about podcasts. In the gist of it, they announced that they would begin to serve podcasts in Google search, allowing users to play audio right there on the website. Google I/O Announcement This is big news for podcasts due to the massive traffic that comes to news from search. Audio Articles could now start to compete with text articles within search. And this is great news for any website publishing audio articles! Publishing audio articles can now help you make gains with Google’s algorithm. Much like how posting content on syndicated sites can push your results to the top of the page — podcasts can now influence Google rankings. The more rich content associated with your domain the better! This is another subtle step towards the integration of voice into every corner of our consumer lives. And like all things SEO it’s imperative to get ahead of the curve NOW! Tips for Audio SEO Optimisation 1. Sign up to SpeechKit. SpeechKit allows publishers to instantly create podcasts of their text existing articles. Audio editions of articles are published as podcasts using our WordPress plugin, RSS or API integrations. Sign up to get started or get in touch at hello@speechkit.io to learn more. 2. Write SEO friendly podcast descriptions. In SpeechKit’s distribution panel you can set up a podcast feed. From the distribution tab select podcast feed and fill out the form. You’ll want to make sure that your title uses keywords that you associate with your site’s content. This too goes for the podcast description. There are many tools out there to help find the best keywords — one we use is SEMrush. 3. Get your podcast live on the stores. SpeechKit provides a unique podcast RSS feed for every podcast that meets the requirements of all popular podcast stores. You can now take this feed and create accounts with Apple Podcast, Spotify, Google Play or wherever else you choose. Google Play for podcasts is only available in selected countries. Yet, this will not affect your podcasts ability to get ranked on the search engine. You’ve now optimised for audio SEO! Now that your podcast is out there, do whatever you can to get it listened to. Repost podcast episodes on social media, embed it into your website, have others post it too. The popularity of the podcast will help your search engine ranking. And remember to stick a SpeechKit player atop each article! More about SpeechKit At SpeechKit we’re helping 100’s of news publishers to automate audio versions of their news articles. Sign up for a free trial and instantly start creating audio for the next generation of news consumer.
21 August 2019. Please note that this post was published before SpeechKit rebranded to BeyondWords. In 2018, 24.8 million people in the UK (45.5%) listened to online audio. Our belief is that this number will continue to grow and with it the percentage of newsreaders who will listen to audio narratives, such as news articles. SpeechKit was designed to help news publishers pivot-to-audio, via audio versions of their news stories, without the time and cost required to narrate them — providing newsreaders with the choice of listening to news articles when reading is undesirable, boosting engagement with audio-streaming native demographics. To keep costs low and the audio scalable we’re using the newly released, neural voices, available through the Amazon text-to-speech (TTS) service (Amazon Polly), to generate lifelike audio. However for news brands seeking to use TTS, like Amazon Polly, to deliver the best audio experience they can, at scale, they’ll need to use SSML (speech-synthesis-markup) tags. Not using SSML will, without a doubt, lead to a sub-par audio experience and dissatisfied listeners. SSML gives you additional control over how Amazon generates speech from text. Enhancing audio with SSML involves inserting specific tags into the text. Doing this manually for a single news article can take time, doing so for all published articles is almost impossible. Amazon Polly supports SSML tags (see table below), but the service does not insert them for you. This usually requires context and Amazon did not develop Amazon Polly just for the news industry. SSML tags supported by SpeechKit We’ve developed a middle-layer, called NewsNet, that, amongst other things, automates the SSML tagging process for news articles using a combination of rule-based and neural-network-based techniques. This post will demonstrate the importance of using SSML when it comes to converting news articles into audio, and highlight the benefit to publishers of using SpeechKit to automate this process. Amazon Polly accepts inputs as either plain text or SSML. For publishers using SpeechKit, NewsNet automatically cleans and converts all plain text from news articles into SSML and encloses SSML tags around paragraphs, sentences, specific words and phrases, amongst a few other things we’ll discuss in another post. , The first step is to wrap the text into a tag. This tells Amazon Polly to process the input as SSML. The second step is to indicate to Amazon Polly that the text should be read as a news item using the tag. The third step, and this is where NewsNet starts to shine, is to tokenize all words, sentences and phrases in the text and apply specific SSML tags to them using either our hardcoded rules or neural nets. , ~~The first of these is the~~
~~and ~~tags that indicate whether a string of text is a sentence or paragraph to ensure that appropriate pauses are inserted into the speech — periods are not always reliable segmentation points in news stories. Other SSML tags inserted using NewsNet include, but are not limited to, , , , and tags.
In some cases, Amazon Polly struggles to pronounce specific words. This is quite common with brands, or in the case below with the president of South Africa. Over time we’ve added hundreds of words, common in the news, and their corrected pronunciation, to NewsNet so that they are detected, tagged, with the phoneme tag, and pronounced correctly appropriately. Cyril Ramaphosa is a South African politician and the fifth and current President of South Africa. Cyril Ramaphosa is a South African politician and the fifth and current President of South Africa.~~ , Different news domains use different acronyms and abbreviations, that when spoken might sound unusual. NewsNet detects...~~
7 August 2019. Please note that this post was published before SpeechKit rebranded to BeyondWords. A news article has always been something that we read and, for the most part, we still do. But what has changed is that our Internet is a lot faster, our phones a lot better, and, coincidentally, we’re a lot busier . So why are we only reading news articles when we have the means to listen to them instead? For many of us, just being able to listen to something is so much easier. You only have to look at the rise in podcast consumption, or audiobook sales over the last few years to conclude that audio is playing an increasingly larger role in our everyday lives. Listen up Listening is convenient because you can get on with other things and potentially learn something new at the same time. We do this with podcasts and audiobooks. Why not news articles? Well-structured news articles, on a topic you’re interested in, can be concise and informative. But they lack the convenience of podcasts and audiobooks. Wouldn’t it be great if we could just listen to them whenever we wanted to? The cost of audio For a news publisher, producing an audio edition of a news article is breathtakingly expensive. It usually requires a voice actor, studio time, some production, and some clever coordination. Producing audio editions for all of them is pretty much impossible. With razor thin margins in the publishing industry, few publishers can afford to invest tens of thousands of dollars in audio despite the demand. Audio needs to change We realised that if news publishers weren’t able to produce narrated audio editions of all their news articles for us to listen to, we’d have to figure out this audio thing for them. That’s when we decided to channel our entrepreneurial spirit and automate the audio process. We decided to build something that reads news articles to us and we called it SpeechKit. Machines and magic News publishers can use SpeechKit to publish audio editions alongside their news articles so that busy (or lazy) folks, like us, don’t have to read them. We use speech synthesis, and a little NLP magic, to compose humanlike audio that, unlike traditional audio editions which use actual humans, cost very little to produce and are available within seconds of an article being published. Just press play We knew that we needed to make it as easy as possible for people to listen to an article, opening another app to do so was out of the question. With this in mind, we decided to get news publishers to do the hard work for you. Now, all you have to do, with publishers that use SpeechKit, is press a play button next to a news article, slip your phone in your pocket and ‘read’ an article — of course, what you choose to listen to is still up to you. You can listen to news articles on BusinessLIVE, Transfer Tavern, Singularity Hub, The Canary, Daily Maverick, and many more. Want to add SpeechKit to your news articles? Contact hello@speechkit.io.