中文 EN FR
Aipunajie Patent Firm

China's First Data IP Registration Certificate Case: Trade Secret Ruling Overturned, Yet RMB 100,000 Still Payable

September 05, 2026 · Data Compliance · Aipunajie Patent Firm / Mili Law Firm

The same Chinese speech dataset: the first-instance court ruled it was a trade secret, while the second-instance court said it was not—but the defendant still had to pay every cent owed: RMB 100,000, plus RMB 2,300 in reasonable expenses.

This is China's first judicial case on the effect of a data IP registration certificate: DataTang (Beijing) v. Yinmou (Shanghai). In September 2021, DataTang launched the "1,505-hour Mandarin Chinese speech data" open-source plan, and in 2023 obtained one of Beijing's first batch of data IP registration certificates. The defendant posted a 200-hour subset on its own official website as a data product for public disclosure and induced registrations and downloads, and was found to have committed unfair competition.

This case clarifies three things, each directly relevant to your AI project.

One Dataset Can Have Three Identities

The court established a three-tier framework of "classified protection, tiered application":

DataTang's speech library falls into the third category—although being public means it cannot be a trade secret, the company's genuine investment means others cannot freeride.

The Registration Certificate Is a "Prima Facie Evidence" Talisman

Absent contrary evidence, the registration certificate = prima facie proof that you enjoy proprietary interests in the data + that the collection was lawfully sourced.

Note the word "prima facie"—it can be rebutted, but the key is: it helps you shift the burden of proof first. If the other side wants to deny it, they must produce their own evidence to overturn it. In litigation, this is a huge first-mover advantage.

Open Source ≠ Waiver of Rights

DataTang used a Creative Commons licence, and the defendant assumed "open source = free commercial use." The court made it clear: whether the open-source licence is followed (especially non-commercial use clauses) is an important consideration in measuring business ethics in the data services field. Unauthorized, freeriding commercial use is unfair competition.

"Register and confirm rights at the input end, keep a documentation trail and maintain confidentiality at the output end. An AI company's data moat is built by actions, not by slogans." — He Zigang | IP Lawyer | Aipunajie · Mili · Najie (20 years of practice, operating three entities with OPC + AI digital employees)

Five Input Compliance Self-Check Lines—Tick Each One Off

① Lawful source — Is the authorization chain for training data complete? Has public data gone through authorized operations? Have you conducted due diligence on data bought from third parties? Is there a documentation trail?

② Prior rights — Does the data contain others' copyrights, trade secrets, or personal information? Article 7 of the Interim Measures for the Administration of Generative AI Services is mandatory: no infringement of others' lawful rights and interests.

③ Open-source licence — Have you read the Creative Commons licence clause by clause for the open-source dataset you cite? Commercial use clauses are red lines; many "free" datasets are actually only "free for non-commercial use."

④ Registration and rights confirmation — For self-built datasets, have you applied for a registration certificate in pilot provinces and cities? As of the end of 2025, over 48,000 certificates had been issued nationwide, with nearly RMB 15 billion in financing credit enhancement. With one certificate, DataTang sold about RMB 95.58 million in data transactions in 2024, up 76% from the previous year.

⑤ Prevent output leaks — This is the most easily overlooked. The State Administration for Market Regulation just announced on August 20, 2026: an algorithm expert in Hangzhou, after leaving, sent out the former employer's AI model-specific prompt templates, review rules, and annotation specifications, and was fined RMB 350,000. The official characterization is firm: natural-language integration solutions and non-standard operating rules can also independently constitute trade secrets.

All Three Layers Must Be Defended: Input → Model → People

In Douyin v. Yiruike "B612 Kaji," China's first case where AI model structure and parameters were protected under the AUCL—the other side copied the structure and parameters of a comic effect model, with a single-case award of RMB 1.6 million, and the Android plus iOS cases totaling about RMB 3.2 million.

From input data (registration certificate), to the intermediate model (trade secret/competitive rights and interests), to departing employees (confidentiality management)—missing one layer costs real money.

Three Suggestions You Can Act On Today

  1. Immediately pull out the list of open-source datasets you are using and check the commercial use clauses of each Creative Commons licence one by one;
  2. Immediately check whether the province or city where your self-built dataset is located has launched a data IP registration pilot; register if you can;
  3. Include prompt templates, model parameters, and review rules in your confidentiality list, sign confidentiality agreements, and manage departures well.

Data compliance: do it early and it is a moat; do it late and it is tuition. Don't wait until you receive a complaint before thinking about reading the open-source licence.

If you've stepped in a pitfall in data compliance, welcome to chat in the comments. Follow the official account "Najie Mili"; next article: When an AI model is copied, how is the RMB 1.6 million award calculated—the dual-track protection playbook for model assets.

This article represents only the author's personal views and does not constitute legal advice. For specific case analysis, welcome to contact us.

This is a machine-translated version of our Chinese original article for reference. The Chinese version is the authoritative source.