Infodemic: Another Facet of Good Old 2020

November 12, 2020

It is difficult to locate non political, non Covid, and non frightening information. I read “Misinformation in the New Normal in a technology publication.” The essay is descriptive; that is, one does not solve a problem or spell out a fix. It’s like a florid passage in James Fennimore Cooper’s novels. There were some factoids in the essay; for example:

According to one piece of research, websites spreading misinformation about the pandemic received nearly half a billion views via Facebook in April alone…

Source? Not stated.

I also noted this statement in the write up:

As defensive measures evolve, so do the attacks, and the further development of deep fake technology is a worrying growth area for misinformation campaigns. Like fake domains, these altered recordings aim to create a veneer of trust in order to seed bad or dangerous information – but deep fakes are now around five years ahead, in technological development terms, of our ability to defend against them.

Five years? That’s another interesting number: 2025. And the lingo like infodemic? Snappy.

I have added the word “infodemic” to my list of interesting neologisms which contain gems like these: neurosymbolic AI, perception hacks, digital detox, and dissonance score.

But the article “Can the Law Stop Internet Bots from Undressing You?” raises another viewpoint about online data; specifically:

For women and men over the age of 18, the production of a sexual pseudo-image of a person is not in itself illegal under international law or in the UK, even if it is produced and distributed without the consent of the person portrayed in the image.

Have government regulators failed? Have educators been unable to impart ethical values to students? Have clever people embraced the methods of some Silicon Valley-type wizards?

Problem solved in 2025?

Stephen E Arnold, November 12, 2020

ThoughtTrace Launches AI Document Comprehension and Management Combo

November 5, 2020

Great idea—Will it work? “ThoughtTrace Unveils the First All-in-One A.I. Document Understanding and Management Platform,” we learn at PR Newswire. The press release explains:

“Today, ThoughtTrace, Inc., the leader in contract and document analytics for asset intensive industries since 2017, announced the official release of their new Document Understanding platform. The new platform combines self-organizing document management with contract analytics and powerful contextual search to discover critical contract data in seconds, condensing weeks of work down to minutes. ‘ThoughtTrace was built to be fundamentally different from both traditional document management and ‘train your own A.I.’ style contract analytics,’ said Nick Vandivere, Chief Executive Officer at ThoughtTrace. ‘With the new platform we are able to completely disrupt traditional approaches to document review that rely on very structured document organization and workflow, and replace that with the ability for the software to actually understand the meaning of the documents being managed. Rather than rigid processes where several different people need to review a document to understand what it says, just ask ThoughtTrace the appropriate question, and it will surface the appropriate results – even for industry specific language, and across thousands to millions of documents.’”

The “appropriate question,” he says. That may be the sticking point for many users. If one can find the magic wording, ThoughtTrace promises to greatly simplify the process of making difficult decisions. We’re told the platform runs on machine learning models tailored to each industry, so no tweaking is required to get started. It can, however, be customized to automate business processes that involve other applications. ThoughtTrace was founded in 1999 and is based in Houston, Texas.

Cynthia Murrell, November 5, 2020

Linear Math Textbook: For Class Room Use or Individual Study

October 30, 2020

Jim Hefferon’s Linear Algebra is a math textbook. You can get it for free by navigating to this page. From Mr. Hefferon’s Web page for the book, you can download a copy and access a range of supplementary materials. These include:

  • Classroom slides
  • Exercise sets
  • A “lab” manual which requires Sage
  • Video.

The book is designed for students who have completed one semester of calculus. Remember: Linear algebra is useful for poking around in search or neutralizing drones. Zaap. Highly recommended.

Stephen E Arnold, October 30, 2020

Content Management: A New Spin

October 27, 2020

What do you get when a young wizard reinvents information management? First, there was records management. Do you know what that was supposed to do? Yep, manage records and know when to destroy them according to applicable guidelines. Next, there was content management. In the era of the Internet, newly minted experts declared that content destined for a Web site had to be management. There were some exciting solutions which made some consultants lots of money; for example, Broadvision/Aurea. Excellent solution. Then there was document management exemplified by companies like Exstream Software which still lives at OpenText as a happy 22 year old solution.) These “disciplines” generated much jargon and handwaving, but most of the chatter sank into data lakes and drowned. Once in a while, like Nessie, an XML/JSON monster emerges and roars, “Success. All your content belong to us.” On the shore of the data lake, eDiscovery vendors shiver in fear. Information management is a scary place.

I read because someone sent me a link, knowing my interest in crazy mid tier consulting speak, to this article: “The Problem with Books of Record and How an EMS Could Help Solve That Problem.” Now here’s the subtitle: “Execution management systems are a new category of software that unlocks value in the hairball of enterprise IT landscapes. Here’s how.”

The acronym EMS means “execution management systems.” Okay. EMS is similar to CMS (content management systems) but with a difference. Execution has a actionable edge. Execution. Get something done. Terminate with extreme prejudice.

Another clarification appears in the write up:

To be a book of record, the data would be in one place, always current and complete. Today’s business systems often have data stored, redundantly, in many places, with many elements incomplete and possibly out of date.

Okay, a book of record and the reference to the existing content chaos which exists in most of these “management” systems.

I am now into new territory. The filing cabinet has yielded to the data lake which suggests dumping everything in one big pool and relying of keywords, Fancy Dan solution like natural language processing, and artificial intelligence to deliver what the person looking for information needs. (The craziness of this approach can be relived by reading about the Google Search Appliance or using an enterprise search system to locate a tweet by a crazed marketer who decided to criticize a competitor after a two hour Zoom meeting followed by a couple of cans of Mountain Dew.)

The write up explains:

Solutions like Celonis’ EMS (execution management) exist because few vendors have focused on all these information handshakes. To create a really efficient business environment, the devil is in the nooks, crannies, handoffs, manual steps, integrations, systems changes, queues, and more. Execution management is about documenting, understanding, integrating, streamlining, optimizing and reengineering how work gets done.  Put simply, Celonis’ tools, in short, document processes, mine what’s happening from the underlying systems to see what kinds of tortured paths are being followed to get work done and then, via benchmarks, best practices and smart automation capabilities, straighten out the flow.

Is this a sales pitch for a company called Celonis?


The firm, according to its Web site, is the number one in the execution management system space. I believe everything I read on the Internet.

Several observations:

  • Automation is a hot topic. Hooking information to workflow makes sense.
  • The word choice or attempt at creating awareness with the EMS moniker could be confusing to some. For me, EMS means emergency management solutions.
  • Founded in 2011, Celonis has ingested (according to Crunchbase) more than $300 million in funding. Investors are optimistic and know that the trajectories of FileNet and FatWire are in their future.

The information management revolution continues. At some point, the problem with information in an organization will be solved. On the other hand, it may be one of those approaching infinity thing-a-ma-bobs. You can’t get there from here.

Some corporate executives experience stress when dealing with content and information challenges: Legal discovery, emails with long forgotten data, and references to documents which no longer “exist.”

Net net: Stress can lead to heart attacks. That’s when the real EMS is needed.

Stephen E Arnold, October 27, 2020

Buzzwords and Baloney: Insecurity Signals? No Way. Do You Like My Hair?

October 22, 2020

People like to sound smart and impressive. The belief is if they appear smart and impressive they will rub shoulders with the best of the best. The Next Web says otherwise in the article: “Using Jargon To Sound Smart? Science Says You’re Just Insecure.”

Apparently people who use too much jargon-use are insecure. Relying on a specialized vocabulary momentarily inflates their ego. This long known truth was proven by the study “Compensatory Conspicuous Communication: Low Status Increases Jargon Use.” The study found that professionals low on the corporate ladder used more acronyms in their written communication and relied on jargon usage when interacting with higher ranks.

All industries have their jargon, but it is alienating to people outside the specific industry. It is even more alienating to others within the industry, because if they are unfamiliar with the term they will not admit it.

Does this mean people on every corporate ladder rung has insecurity? Yup.

Unfortunately you cannot beat jargon users so it is better to join the herd:

“As much as it’s annoying and superfluous, jargon is unlikely to go away. So you literally have two choices: you can embrace it or ignore it. I’m of the opinion that if you can’t beat them, you join them. How? By using a technology bullsh*t generator — yes, you’ve read that correctly. This tool won’t change your life but you’ll definitely have some fun.”

Another fun thing to do with jargon enthusiasts is make up words. It takes practice, but if you speak confidently enough you will soon be “proclaving” [sic] people. Cloudify too.

Whitney Grace, October 23, 2020

Text Analytics: Are These Really the Companies to Watch in the Next 12 Weeks?

October 16, 2020

DarkCyber spotted “Top 10 Text Analytics Companies to Watch in 2020.” Let’s take a quick look at some basic details about each firm:

Alkymi, founded in 2017, makes an email indexing system. The system, according to the company’s Web site, “understands documents using deep learning and visual analysis paired with your human in-the-loop expertise.” Interesting but text analytics appears to be a component of a much larger system. What’s interesting is that the business relies in some degree upon Amazon Web Services. The company’s Web site is

Aylien Ltd., based in Ireland, appears to be a company with text analysis technology. However, the company’s system is used to create intelligence reports for analysts; for example, government intelligence officers, business analysts, and media outlets. Founded in 2010, the company’s Web site is

Hewlett Packard Enterprise. The inclusion of HPE was a bit of a surprise. This outfit once owned the Autonomy technology, but divested itself of the software and services. To replace Autonomy, the company developed “Advanced Text Analysis” which appears to be an enterprise search centric system. The service is available as a Microsoft Azure function and offers 60 APIs (which seems particularly generous) “that deliver deep learning analytics on a wide range of data.” The company’s Web site is One product name jumped out: Ezmeral which maybe a made up word.

InData Labs lists data science, AI, AI driven mobile app development, computer vision, machine learning, data capture and optical character recognition, and big data solutions as its services. Its products include face recognition and natural language processing. Perhaps it is the NLP product which equates to text analytics? The firm’s Web site is The company was founded in 2014 and operates from Belarus and has a San Francisco presence.

Kapiche, founded in 2016, focuses on “customer insights”. Customer feedback yields insight with “no set up, no manual coding, and results you can trust,” according to the company. The text analytics snaps into services like Survey Monkey and Google Forms, among others. Clients include Target and Toyota. The company is based in Australia with an office in Denver, Colorado. The firm’s Web site is The firm offers applied text analytics.

Lexalytics, founded in 2003, was one of the first standalone text analytics vendors. The company’s system allows customers to “tell powerful stories from complex text data.” DarkCyber prefers to learn “stories” from the data, however. In the last 17 years, the company has not gone public nor been acquired. The firm’s Web site is

MindGap. The MindGap identified in the article is in the business of providing “AI for business.” the company appears to be a mash up of artificial intelligence and “top tier strategy consulting:. That may be true, but we did not spot text analytics among the core competencies. The firm’s clients include, Gazprom, Yandex, and Huawei. The firm’s Web site is The firm lists two employees on LinkedIn.

Primer has ingested about $60 million in venture funding since it was founded  in 2015. The company ingests text and outputs reports. The company was founded by the individual who set up Quid, another analytics company. Government and business analysts consume the outputs of the Primer system. The company’s Web site is

Semeon Analytics, now a unit of Datametrex, provides “custom language and sentiment ontology” services. Indexing and entity extraction, among other NLP modules, allows the system to deliver “insight analysis, rapid insights, and sentiment of the highest precision on the market today.” The Semeon Web site is still online at

ThoughtTrace appears to focus on analysis of text in contracts. The firm’s Web site says that its software can “find critical contract facts and opportunities.” Text analytics? Possibly, but the wording suggests search and retrieval. The company has a focus on oil and gas and other verticals. The firm’s Web site is (Note that the design of the Web site creates some challenges for a person looking for information.) The company, according to Crunchbase, was founded in 1999, and has three employees.

Three companies are what DarkCyber would consider text analytics firms: Aylien, Lexalytics, and Primer. The other firms mash up artificial intelligence, machine learning, and text analytics to deliver solutions which are essentially indexing and workflow tools.

Other observations include:

  1. The list is not a reliable place to locate flagship vendors; specifically, only three of the 10 companies cited in the article could be considered contenders in this sector.
  2. The text analytics capabilities and applications are scattered. A person looking for a system which is designed to handle email would have to examine the 10 listings and work from a single pointer, Alkymi.
  3. The selection of vendors confuses technical disciplines; for example, AI, machine learning, NLP, etc.

The list appears to have been generated in a short Zoom meeting, not via a rigorous selection and analysis process. Perhaps one of the vendors’ text analytics systems could have been used. Primer’s system comes to mind as one possibility. But that, of course, is work for a real journalist today.

Stephen E Arnold, October 16, 2020

Sentiment Analysis with Feeling

September 25, 2020

As AI technology progresses, so too does the field of sentiment analysis. What could go wrong? Sinapticas explores “How Algorithms Discern Our Mood from What We Write Online.” Reporter Dana Mackenzie begins with an example we can truly relate to right now:

“Many people have declared 2020 the worst year ever. While such a description may seem hopelessly subjective, according to one measure, it’s true. That yardstick is the Hedonometer, a computerized way of assessing both our happiness and our despair. It runs day in and day out on computers at the University of Vermont (UVM), where it scrapes some 50 million tweets per day off Twitter and then gives a quick-and-dirty read of the public’s mood. According to the Hedonometer, 2020 has been by far the most horrible year since it began keeping track in 2008. The Hedonometer is a relatively recent incarnation of a task computer scientists have been working on for more than 50 years: using computers to assess words’ emotional tone. To build the Hedonometer, UVM computer scientist Chris Danforth had to teach a machine to understand the emotions behind those tweets — no human could possibly read them all. This process, called sentiment analysis, has made major advances in recent years and is finding more and more uses.”

The accompanying “Average Happiness for Twitter” graph is worth a gander, and well illustrates the concept (and the ride that has been 2020 thus far). The article is a good introduction to sentiment analysis. It contrasts lexicon-based with the more complex neural network approaches. We learn neural networks may never completely eclipse lexicon-based systems because of the immense computing power required for the latter. Hedonometer, for example, uses a lexicon.

Mackenzie also describes several applications of sentiment analysis, like predicting mental health, assessing prevailing attitudes on issues of the day, and, of course, supplying business intelligence. And the hedonism of the hedonometer, of course.

Cynthia Murrell, September 25, 2020

Fixing Language: No Problem

August 7, 2020

Many years ago I studied with a fellow who was the world’s expert on the morpheme _burger. Yep, hamburger, cheeseburger, dumbburger, nothingburger, and so on. Dr. Lev Sudek (I think that was his last name but after 50 years former teachers blur in my mind like a smidgen of mustard on a stupidburger.) I do recall his lecture on Indo-European languages, the importance of Sanskrit, and the complexity of Lithuanian nouns. (Why Lithuanian? Many, many inflections.) Those languages evolving or de-volving from Sanskrit or ur-Sanskrit differentiated among male, female, singular, neuter, plural, and others. I am thinking 16 for nouns but again I am blurring the Sriacha on the Incredible burger.

This morning, as I wandered past the Memoryburger Restaurant, I spotted “These Are the Most Gender-Biased Languages in the World (Hint: English Has a Problem).” The write up points out that Carnegie Mellon analyzed languages and created a list of biased languages. What are the languages with an implicit problem regarding bias? Here a list of the top 10 gender abusing, sexist pig languages:

  1. Danish
  2. German
  3. Norwegian
  4. Dutch
  5. Romanian
  6. English
  7. Hebrew
  8. Swedish
  9. Mandarin
  10. Persian

English is number 6, and if I understand Fast Company’s headline, English has a problem. Apparently Chinese and Persian do too, but the write up tiptoes around these linguistic land mines. Go with the Covid ridden, socially unstable, and financially stressed English speakers. Yes, ignore the Danes, the Germans, Norwegians, Dutch, and Romanians.

So what’s the fix for the offensive English speakers? The write up dodges this question, narrowing to algorithmic bias. I learned:

The implications are profound: This may partially explain where some early stereotypes about gender and work come from. Children as young as 2 exercise these biases, which cannot be explained by kids’ lived experiences (such as their own parents’ jobs, or seeing, say, many female nurses). The results could also be useful in combating algorithmic bias.

Profound indeed. But the French have a simple, logical, and  “c’est top” solution. The Académie Française. This outfit is the reason why an American draws a sneer when asking where the computer store is in Nimes. The Académie Française does not want anyone trying to speak French to use a disgraced term like computer.

How’s that working out? Hashtag and Franglish are chugging right along. That means that legislating language is not getting much traction. You can read a 290 page dissertation about the dust up. Check out “The Non Sexist Language Debate in French and English.” A real thriller.

The likelihood of enforcing specific language and usage changes on the 10 worst offenders strikes me as slim. Language changes, and I am not sure the morpheme –burger expert understood decades ago how politicallycorrectburgers could fit into an intellectual menu.

Stephen E Arnold, August 7, 2020

Tick Tock Becomes Tit for Tat: The Apple and Xiao-i Issue

August 5, 2020

Okay, let’s get the company names out of the way:

  • Shanghai Zhizhen Network Technology Company is known as Zhizhen
  • Zhizhen is also known as Xiao-i
  • Apple is the outfit with the virtual assistant Siri.

Zhizhen owns a patent for a virtual assistant. In 2013, Apple was sued for violating a Chinese patent. Apple let loose a flock of legal eagles to demonstrate that its patents were in force and that a Chinese voice recognition patent was invalid. The Chinese court denied Apple’s argument.

Tick tock tick tock went the clock. Then the alarm sounded. Xiao-i owns the Chinese patent, and that entity is suing Apple.

Apple Faces $1.4B Suit from Chinese AI Company” reports:

Shanghai Zhizhen Network Technology Co. said in a statement on Monday it was suing Apple for an estimated 10 billion yuan ($1.43 billion) in damages in a Shanghai court, alleging the iPhone and iPad maker’s products violated a patent the Chinese company owns for a virtual assistant whose technical architecture is similar to Siri. Siri, a voice-activated function in Apple’s smartphones and laptops, allows users to dictate text messages or set alarms on their devices.

But more than the money, the Xiao-i outfit “asked Apple to stop sales, production, and the use of products fluting such a patent.”

Coincidence? Maybe. The US wants to curtail TikTok, and now Xiao-i wants to put a crimp in Apple’s China revenues.

Several observations:

  • More trade related issues are likely
  • Intellectual property disputes will become more frequent. China will use its patents to inhibit American business. This is a glimpse of a future in which the loss of American knowledge value will add friction to the US activities
  • Downstream consequences are likely to ripple through non-Chinese suppliers of components and services to Apple. China is using Apple to make a point about the value of Chinese intellectual property and the influence of today’s China.

Just as China has asserted is cyber capabilities, the Apple patent dispute — regardless of its outcome — is another example of China’s understanding of American tactics, modifying them, and using them to try to gain increased economic, technical, and financial advantage.

Stephen E Arnold, August 3, 2020

Natural Language Processing: Useful Papers Selected by an Informed Human

July 28, 2020

Nope, no artificial intelligence involved in this curated list of papers from a recent natural language conference. Ten papers are available with a mouse click. Quick takeaway: Adversarial methods seem to be a hot ticket. Navigate to “The Ten Must Read NLP/NLU Papers from the ICLR 2020 Conference.” Useful editorial effort and a clear, adult presentation of the bibliographic information. Kudos to jakubczakon.

Stephen E Arnold, July 27, 2020

Next Page »

  • Archives

  • Recent Posts

  • Meta